Visual media search method, and device and storage medium

By allowing users to use search statements containing time search ranges and keywords, the system can recall photos and videos whose acquisition time is outside the range but matches the keywords, solving the problem that the same series of photos or videos cannot be recalled in the prior art, and improving the user's search experience.

WO2025108044A1PCT designated stage expired Publication Date: 2025-05-30HONOR DEVICE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/129183
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-21
Filing Date
2024-10-31
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to recall photos and videos taken by users outside a specific time range but matched the search keywords through simple search statements, resulting in the same series of photos or videos taken by users before and after the time search range cannot be recalled, affecting the user's search experience.

Method used

Provides a visual media search method, allowing users to search photos and videos through search statements that include both the time search range and keywords. The system will recall photos and videos whose acquisition time is outside the time search range but match the keywords.

Benefits of technology

Ensure that the same series of photos or videos taken by users before and after the time search range can be recalled, improving the user's search experience and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024129183_30052025_PF_FP_ABST
    Figure CN2024129183_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a visual media search method, and a device and a storage medium. The method is adapted to an electronic device, and comprises: displaying a first interface, wherein the first interface comprises a search box; receiving a search operation in respect of a search statement which is input via the search box, wherein the search statement comprises a time search range and a first keyword, and the starting time of the time search range is a first time point, and the ending time thereof is a second time point; and displaying a first search result, the first search result corresponding to a first visual media file, wherein the first visual media file matches the first keyword, and the collection time of the first visual media file is before the first time point or after the second time point. The technical solution provided in the embodiments of the present application can ensure that the same series of photographs or videos captured by a user before and after a time search range which is input by the user can be recalled, thereby improving the user search experience.
Need to check novelty before this filing date? Find Prior Art

Description

Visual media search method, device and storage medium

[0001] Cross-references

[0002] This application refers to Chinese Patent Application No. 2023115620330 filed on November 21, 2023, entitled “Visual Media Search Method, Device and Storage Medium” and Chinese Patent Application No. 2023115640118 filed on November 21, 2023, entitled “Visual Media Search Method, Device and Storage Medium”, both of which are incorporated herein by reference in their entirety. Technical Field

[0003] The present application relates to the field of terminal technology, and in particular to a visual media search method, device, and storage medium. Background Art

[0004] With the popularity of smart devices, more and more users are using them, such as mobile phones, to take photos and videos, and store them in the gallery of their electronic devices, thereby recording every detail of their lives. Furthermore, users can also download images and take screenshots of their mobile phone interfaces, and store these downloaded images and screenshots in the gallery of their electronic devices.

[0005] To facilitate users in managing and viewing pictures in the terminal, the terminal device's gallery application or other similar applications are equipped with picture management and picture search functions. For example, the gallery application in the terminal can classify the pictures in the terminal into corresponding albums based on information such as the time and place where the pictures were taken. Users can view related pictures by searching for information such as time and place.

[0006] Summary of the Invention

[0007] Multiple aspects of the present application provide a visual media search method, device and storage medium, which can enable users to search for photos and videos through search statements that simultaneously include a time search range and keywords, and can search for photos and videos whose acquisition time is outside the time search range but matches the keyword, so as to ensure that the same series of photos or videos taken by the user before and after the time search range can be recalled, thereby improving the user's search experience.

[0008] In a first aspect, a visual media search method is provided, applicable to an electronic device, comprising:

[0009] Displaying a first interface; the first interface includes a search box;

[0010] Receiving a search operation of a search statement input into the search box; wherein the search statement includes a time search range and a first keyword, and the start time of the time search range is a first time point and the end time is a second time point;

[0011] displaying a first search result, the first search result corresponding to a first visual media file;

[0012] wherein the first visual media file matches the first keyword;

[0013] The acquisition time of the first visual media file is before the first time point or after the second time point.

[0014] It can be seen that in the technical solution provided by the present application, when a user uses a search statement that includes both a time search range and keywords to search for pictures and videos, not all photos and videos taken by the user before the start time point or after the end time point of the time search range will be filtered out. Instead, photos and videos taken by the user before the start time point or after the end time point of the time search range that match the first keyword will be recalled, thereby improving the user's search experience.

[0015] Optionally, when the acquisition time of the first visual media file is before the first time point, the time difference between the acquisition time of the first visual media file and the first time point is less than a second threshold.

[0016] Optionally, when the acquisition time of the first visual media file is after the second time point, the time difference between the acquisition time of the first visual media file and the second time point is less than a third threshold.

[0017] That is, it is possible to recall photos and videos that match the first keyword and are taken by the user within a limited time period before the start time point or within a limited time period after the end time point of the time search range.

[0018] The above-mentioned second threshold and the second threshold may be equal or unequal in size, and the specific value may be set according to actual needs, and this application does not make any specific limitation on this.

[0019] In one possible implementation, after receiving a search operation of a search statement input into the search box, the method includes: displaying a second search result, where the second search result corresponds to a second visual media file;

[0020] The visual content of the second visual media file matches the first keyword;

[0021] The acquisition time of the second visual media file is between the first time point and the second time point, and the difference between the acquisition time of the first visual media file and the acquisition time of the second visual media file is less than a first threshold.

[0022] That is to say, this solution can not only recall photos and videos that match the first keyword taken by the user between the start time point and the end time of the time search range, but also recall photos and videos that match the first keyword taken by the user within a limited time period before the start time point or within a limited time period after the end time of the time search range.

[0023] In one possible implementation, displaying the first interface includes:

[0024] In response to an operation of opening the gallery application, displaying the first interface;

[0025] or, in response to an operation of opening the negative one screen triggered on the main screen of the electronic device, displaying the first interface;

[0026] Or, in response to a pull-down search operation triggered on the home screen of the electronic device, the first interface is displayed.

[0027] That is, users can search for photos, videos and other media through the gallery application interface, the negative one screen interface or the pull-down search interface.

[0028] In one possible implementation, displaying the first search result includes:

[0029] A first picture is displayed, where the first picture is a proportional thumbnail of the first visual media file or a proportional thumbnail of any image frame.

[0030] When the first visual media file is a picture, the first picture is a proportional thumbnail of the picture; when the first visual media file is a video, the first picture is a proportional thumbnail of any image frame (eg, the first frame) of the video.

[0031] It can be understood as displaying thumbnails of search results through a grid page.

[0032] In one possible implementation, the method further includes: displaying a third image; the third image being a proportional thumbnail of a third visual media file or a proportional thumbnail of any image frame; the third visual media file matching the first keyword;

[0033] The first picture and the third picture are displayed on the second interface; the first picture is displayed in front of the third picture; the matching degree between the visual media file corresponding to the first picture and the first keyword is a first matching degree; the matching degree between the visual media file corresponding to the second picture and the first keyword is a second matching degree; the first matching degree is greater than the second matching degree.

[0034] It can be understood that the thumbnails of the multiple visual media files obtained by the search are displayed in order on the grid page according to the matching degree between the respective visual media files and the first keyword from high to low, so as to facilitate users to quickly find the visual media they want.

[0035] In one possible implementation, the visual content of the first visual media file matches the first keyword; the visual content is data that needs to be acquired through a natural image understanding model.

[0036] In one possible implementation, the first keyword is matched with the first visual media file by comparing images and texts in a pre-trained CLIP model.

[0037] Through model training, the CLIP model can map visual media and text into a unified vector space to fully understand the relationship between different modal data in text and vision, which can improve matching accuracy and thus improve recall accuracy.

[0038] In a second aspect, the present application provides a visual media search method applicable to an electronic device, comprising:

[0039] Displaying a first interface; the first interface includes a search box;

[0040] receiving a search operation of a search statement input into the search box, wherein the search statement includes a time search range;

[0041] Using a clustering algorithm, clustering the plurality of first visual media files according to their acquisition time to obtain K clusters, wherein the K clusters have different acquisition time ranges; wherein K ≥ 1 and is an integer;

[0042] Determining a target cluster according to whether the acquisition time range partially overlaps with the time search range;

[0043] Determining search results of the search statement according to the target cluster;

[0044] The search results are displayed.

[0045] In this solution, visual media files collected at relatively close times are clustered into a cluster using a clustering algorithm. That is, the collection times of the visual media files in each cluster are relatively close.

[0046] A cluster whose acquisition time range partially overlaps with the time search range means that the cluster includes visual media whose acquisition time is before the start time point of the time search range or visual media whose acquisition time is after the end time point of the time search range. Moreover, the acquisition time of the visual media in the cluster whose acquisition time is before the start time point of the time search range or after the end time point of the time search range is relatively close to the acquisition time of other visual media in the cluster. It can be seen that this solution can recall the same series of photos or videos that the user took continuously before and after the start time point or end time point of the time search range, so as to improve the user's search experience.

[0047] In one possible implementation, the start time of the acquisition time range of each cluster is the acquisition time of the earliest acquired visual media file in the cluster, and the end time is the acquisition time of the latest acquired visual media file in the cluster; there is no intersection between the acquisition time ranges of the K clusters.

[0048] In one possible implementation, the start time of the time search range is a first time point, and the end time is a second time point;

[0049] Determining a target cluster according to whether the acquisition time range partially overlaps with the time search range includes:

[0050] When the start time of the acquisition time range of the kth cluster among the K clusters is between the first time point and the second time point and the end time is not between the first time point and the second time point, determining that the acquisition time range of the kth cluster partially overlaps with the time search range;

[0051] When the end time of the acquisition time range of the k-th cluster is between the first time point and the second time point and the start time is not between the first time point and the second time point, determining that the acquisition time range of the k-th cluster partially overlaps with the time search range;

[0052] When the start time of the acquisition time range of the k-th cluster is less than the first time point and the end time is greater than the second time point, determining that the acquisition time range of the k-th cluster partially overlaps with the time search range; k is an integer, 1≤k≤K;

[0053] A cluster whose acquisition time range partially overlaps with the time search range among the K clusters is determined as a target cluster.

[0054] Among them, the starting time is between the first time point and the second time point, which means that the starting time is greater than or equal to the first time point and less than or equal to the second time point; the starting time is not between the first time point and the second time point, which means that the starting time is less than the first time point or greater than the second time point; the ending time is between the first time point and the second time point, which means that the ending time is greater than or equal to the first time point and less than or equal to the second time point; the ending time is not between the first time point and the second time point, which means that the ending time is less than the first time point or greater than the second time point.

[0055] In one possible implementation, the search results include a plurality of target clusters; and the method further includes:

[0056] Determining a display order of the plurality of target clusters according to a chronological order of start times of the acquisition time ranges of the plurality of target clusters;

[0057] Display the search results, including:

[0058] The plurality of target clusters are displayed in a display order of the plurality of target clusters.

[0059] It is understandable that, on the grid page, multiple target clusters are displayed in sequence according to the start time of the collection time range of the multiple target clusters. The collection time range of the target cluster displayed earlier on the grid page starts earlier than the collection time range of the target cluster displayed later.

[0060] In one possible implementation, the method further includes:

[0061] Performing semantic understanding on the search sentence to obtain a semantic vector of the first sentence;

[0062] Obtaining visual semantic vectors of multiple visual media files; the visual semantic vector of each visual media file is obtained by semantically understanding an image or image frame of the visual media file using a natural image understanding model;

[0063] Determining, from the plurality of visual media files, M candidate visual media files whose visual semantic vectors match the first sentence semantic vector; wherein M is an integer greater than 1;

[0064] The plurality of first visual media files are determined based on the M candidate visual media files.

[0065] In other words, the multiple first visual media files are retrieved by matching the visual semantic vectors of the visual media files with the sentence semantic vectors of the search phrase (hereinafter referred to as the first branch of retrieval). This ensures that the visual content of the visual media in the final search results semantically matches the search phrase.

[0066] Specifically, the CLIP model can be used to match visual media files with search statements.

[0067] In one possible implementation, a clustering algorithm is used to cluster the plurality of first visual media files according to their acquisition time to obtain K clusters, including:

[0068] If the proportion of semantics related to the visual content in the search statement is greater than or equal to a preset proportion threshold, clustering the multiple first visual media files according to their acquisition time using a clustering algorithm to obtain K clusters; the visual content is data that needs to be acquired by a natural image understanding model;

[0069] The method further comprises:

[0070] obtaining properties of the plurality of visual media files;

[0071] Determining, from the plurality of visual media files, a plurality of second visual media files whose attributes match the semantic subject in the search statement (hereinafter referred to as second branch recall);

[0072] If the proportion of semantics related to visual content in the search statement is less than the preset proportion threshold, the search results of the search statement are determined based on the multiple second visual media.

[0073] That is, when the proportion of semantics related to visual content in the search statement is greater than or equal to a preset proportion threshold (ie, a high visual semantic search statement), the visual media recalled by the first branch is used as the candidate set.

[0074] When the proportion of semantics related to visual content in the search statement is less than a preset proportion threshold (i.e., a low visual semantic search statement), the visual media recalled by the second branch is used as the candidate set.

[0075] That is, for search statements with high visual semantics, the visual semantics of the visual media are used to recall and obtain candidate sets; for search statements with low visual semantics, the attributes of the visual media (such as time and place) are used to recall and obtain candidate sets. This can improve search accuracy.

[0076] In this solution, for high visual semantic search statements and low visual semantic search statements, the candidate sets are recalled using their respective adapted recall methods, thereby improving the completeness and accuracy of the search results and avoiding the incompleteness or inaccuracy of the search results caused by using the same recall method.

[0077] For example, when a search statement is low in visual semantics (e.g., photos taken in Beijing this year), if the recall is based on the matching of the search statement with the visual content of multiple visual media, many photos taken by users in Beijing this year will be filtered out, resulting in incomplete search results. When a search statement is high in visual semantics (e.g., the sky taken in Beijing this year), if the recall is based on the matching of the time or location in the search statement with the attributes of multiple visual media, many photos unrelated to "sky" will be recalled, resulting in inaccurate search results.

[0078] In one possible implementation, the method further includes:

[0079] removing semantic entities irrelevant to the visual content from the search statement to obtain a rewritten search statement;

[0080] Performing semantic understanding on the rewritten search sentence to obtain a second sentence semantic vector;

[0081] Determining, from the plurality of visual media files, a plurality of reference visual media files whose visual semantic vectors match the second sentence semantic vector;

[0082] determining first representative visual semantic vectors corresponding to the plurality of first visual media files and second representative visual semantic vectors corresponding to the plurality of reference visual media;

[0083] determining a semantic proportion related to the visual content in the search statement based on a vector similarity between the first difference vector and the second difference vector;

[0084] The first difference vector represents the difference between the first sentence semantic vector and the first representative visual semantic vector; the second difference vector represents the difference between the second sentence semantic vector and the second representative visual semantic vector.

[0085] The above-mentioned semantic subjects that are not related to the visual content mainly refer to semantic subjects related to time (for example, this year) and semantic subjects related to place (for example, Shanghai). These two types of semantic subjects can be identified by named entity recognition (NER) technology. Therefore, removing the semantic subjects that are not related to the visual content in the search statement to obtain a rewritten search statement can specifically include: removing the semantic subjects related to time and / or the semantic subjects related to place in the search statement to obtain a rewritten search statement.

[0086] The greater the vector similarity between the first and second difference vectors, the greater the proportion of semantics related to the visual content in the search statement. The smaller the vector similarity between the first and second difference vectors, the smaller the proportion of semantics related to the visual content in the search statement. In other words, the proportion of semantics related to the visual content in the search statement is positively correlated with the vector similarity between the first and second difference vectors.

[0087] In one possible implementation, determining the plurality of first visual media files based on the M candidate visual media files includes:

[0088] Determining N dimensions to be matched corresponding to the search statement based on a semantic subject related to the visual content in the search statement, where N is an integer greater than 1;

[0089] Obtaining weights of the N dimensions to be matched; wherein the weight of the jth dimension to be matched is positively correlated with the degree of variation in the matching degree of the visual contents of the M candidate visual media files on the jth dimension to be matched; j is an integer, and the value of j ranges from 1 to N;

[0090] According to the weights of the N dimensions to be matched, weighted summation is performed on the matching degrees of the visual content of the i-th candidate visual media file in each of the N dimensions to be matched, to obtain a comprehensive matching degree of the i-th candidate visual media file; i is an integer, and the value of i ranges from 1 to M;

[0091] The plurality of first visual media files whose comprehensive matching degrees meet preset requirements are determined from the M candidate visual media files.

[0092] The M candidate visual media files are obtained by matching the visual semantic vectors of the visual media files with the sentence semantic vectors of the search statement. That is, when recalling the M candidate visual media files, the degree of match between the visual semantics of the visual media files and the user's entire search statement is considered, without considering the degree of match between the visual semantics of the visual media files and the semantic subject of the search statement that the user visually focused on, resulting in the recall of some inaccurate photos and videos. Therefore, in this solution, the results of the first branch recall are fine-tuned by using the degree of match between the visual semantics of the visual media and the "semantic subject related to the visual content" to filter out inaccurate visual media.

[0093] Specifically, the semantic entities related to the visual content in the search statement and / or the search statement itself can be used as different dimensions to be matched. Different dimensions to be matched play different roles in fine-tuning, that is, the weights of different dimensions to be matched are different. The weight of each dimension to be matched is determined based on the degree of variation in the matching degree of the visual content of the M candidate visual media files on the dimension to be matched. The greater the degree of variation, the greater the role played by the dimension to be matched during fine-tuning, and therefore, the greater its weight. In this way, some inaccurate visual media such as photos and videos can be effectively excluded.

[0094] In one possible implementation, the method further includes:

[0095] Matching a search term in the search statement with a plurality of preset tags to determine whether the search term belongs to a semantic subject associated with the tag; the tag is used to describe visual content;

[0096] The semantic subject related to the tag in the search sentence is determined as the semantic subject related to the visual content in the search sentence.

[0097] In other words, a search term that matches one of the preset tags belongs to the semantic subject associated with the tag. Since tags are related to visual content, the semantic subject associated with the tag also belongs to the semantic subject associated with the visual content. In this solution, by presetting multiple tags, it is relatively easy to determine the semantic subject that the user is most visually interested in from the search statement.

[0098] In one possible implementation, the clustering algorithm includes: a density-based clustering algorithm;

[0099] The method further comprises:

[0100] According to the time search range, a first parameter involved in the density-based clustering algorithm is determined; the first parameter is used to describe the neighborhood radius of the data point; the first parameter is positively correlated with the duration corresponding to the time search range.

[0101] The method may further include:

[0102] determining the number of visual media files in the plurality of first visual media files whose acquisition time falls within the time search range;

[0103] A second parameter involved in the density-based clustering algorithm is determined based on the ratio of the number of visual media files to the total duration corresponding to the time search range; the second parameter is used to describe the minimum number of data points in the neighborhood of a data point.

[0104] In this embodiment, the first parameter and the second parameter involved in the density-based clustering algorithm are dynamically adjusted according to the time search range in the user's search statement to adapt to different search scenarios, thereby improving the user's search experience.

[0105] In one possible implementation, determining search results of the search statement according to the target cluster includes:

[0106] If the search statement includes a semantic subject related to a location, filtering the visual media files in the target cluster according to the semantic subject related to the location and the collection location attributes of the visual media files in the target cluster;

[0107] If the search statement includes a semantic subject related to character relationships, filtering the visual media files in the target cluster according to the semantic subject related to character relationships and the character relationship attributes of the visual media files in the target cluster; and / or

[0108] If the search statement includes a semantic subject related to a person's name, the visual media files in the target cluster are filtered according to the semantic subject related to the person's name and the person's name attribute of the visual media files in the target cluster.

[0109] In this solution, the recalled visual media are filtered by location, person relationship, and name.

[0110] In a third aspect, a visual media search method is provided, applicable to an electronic device, comprising:

[0111] Displaying a first interface; the first interface includes a search box;

[0112] receiving a search operation of a search statement input into the search box;

[0113] If the proportion of semantics related to visual content in the search statement is greater than or equal to a preset proportion threshold, determining search results for the search statement based on a plurality of first visual media files; the plurality of first visual media files are determined from the plurality of visual media based on a match between the search statement and the visual content of the plurality of visual media; the visual content being data that needs to be acquired through a natural image understanding model;

[0114] If the proportion of semantics related to the visual content in the search statement is less than the preset proportion threshold, determining search results for the search statement based on a plurality of second visual media; the plurality of second visual media files are determined from the plurality of visual media based on a match between semantic entities in the search statement that are not related to the visual content and attributes of the plurality of visual media;

[0115] The search results are displayed.

[0116] Specifically, the matching of the search statement with the visual content of multiple visual media may include: the matching of the first sentence semantic vector of the search statement with the visual semantic vectors of multiple visual media (for example, matching degree) or the matching of the subject semantic vector of the semantic subject related to the visual content in the search statement with the visual semantic vectors of multiple visual media (for example, matching degree).

[0117] In this solution, a search statement with a visual content-related semantics ratio greater than or equal to a preset threshold is considered a high-visual-semantic search statement; a search statement with a visual content-related semantics ratio less than the preset threshold is considered a low-visual-semantic search statement. For high-visual-semantic and low-visual-semantic search statements, candidate sets are recalled using their respective adaptive recall methods, thereby improving the completeness and accuracy of search results and avoiding the incompleteness or inaccuracy of search results caused by using the same recall method.

[0118] For example, when a search statement is low in visual semantics (e.g., photos taken in Beijing this year), if the recall is based on the matching of the search statement with the visual content of multiple visual media, many photos taken by users in Beijing this year will be filtered out, resulting in incomplete search results. When a search statement is high in visual semantics (e.g., the sky taken in Beijing this year), if the recall is based on the matching of the time or location in the search statement with the attributes of multiple visual media, many photos unrelated to "sky" will be recalled, resulting in inaccurate search results.

[0119] In one possible implementation, the above method further includes:

[0120] removing semantic entities irrelevant to the visual content from the search statement to obtain a rewritten search statement;

[0121] determining a plurality of reference visual media from the plurality of visual media based on a matching condition between the rewritten search statement and the visual contents of the plurality of visual media;

[0122] determining first representative visual semantic vectors corresponding to the plurality of first visual media files and second representative visual semantic vectors corresponding to the plurality of reference visual media;

[0123] determining a semantic proportion related to the visual content in the search statement based on a vector similarity between the first difference vector and the second difference vector;

[0124] The first difference vector represents the difference between the first sentence semantic vector and the first representative visual semantic vector; the second difference vector represents the difference between the second sentence semantic vector and the second representative visual semantic vector.

[0125] The greater the vector similarity between the first and second difference vectors, the greater the proportion of semantics related to the visual content in the search statement. The smaller the vector similarity between the first and second difference vectors, the smaller the proportion of semantics related to the visual content in the search statement. In other words, the proportion of semantics related to the visual content in the search statement is positively correlated with the vector similarity between the first and second difference vectors.

[0126] In practice, the structure of a sentence is relatively complex, making it difficult to simply calculate the proportion of semantics related to the visual content in the sentence based on the number of semantic entities related to the visual content and the number of semantic entities unrelated to the visual content. The above method leverages the word analogy properties of distributed representation vectors to accurately determine the proportion of semantics related to the visual content in the search sentence.

[0127] In one possible implementation, removing semantic entities irrelevant to the visual content from the search statement to obtain a rewritten search statement includes:

[0128] The time-related semantic subject and / or the location-related semantic subject in the search statement are removed to obtain a rewritten search statement.

[0129] In one possible implementation, the method further includes:

[0130] Performing semantic understanding on the search sentence to obtain a semantic vector of the first sentence;

[0131] Obtaining visual semantic vectors of the plurality of visual media files; the visual semantic vector of each visual media file is obtained by semantically understanding an image or image frame of the visual media file using a natural image understanding model;

[0132] Determining, from the plurality of visual media files, M candidate visual media files whose visual semantic vectors match the first sentence semantic vector; wherein M is an integer greater than 1;

[0133] The plurality of first visual media files are determined based on the M candidate visual media files.

[0134] In other words, the multiple first visual media files are retrieved by matching the visual semantic vectors of the visual media files with the sentence semantic vectors of the search phrase (hereinafter referred to as the first branch of retrieval). This ensures that the visual content of the visual media in the final search results semantically matches the search phrase.

[0135] In one possible implementation, determining the plurality of first visual media files based on the M candidate visual media files includes:

[0136] Determining N dimensions to be matched corresponding to the search statement based on a semantic subject related to the visual content in the search statement, where N is an integer greater than 1;

[0137] Obtaining weights of the N dimensions to be matched; wherein the weight of the jth dimension to be matched is positively correlated with the degree of variation in the matching degree of the visual contents of the M candidate visual media files on the jth dimension to be matched; j is an integer, and the value of j ranges from 1 to N;

[0138] According to the weights of the N dimensions to be matched, weighted summation is performed on the matching degrees of the visual content of the i-th candidate visual media file in each of the N dimensions to be matched, to obtain a comprehensive matching degree of the i-th candidate visual media file; i is an integer, and the value of i ranges from 1 to M;

[0139] The plurality of first visual media files whose comprehensive matching degrees meet preset requirements are determined from the M candidate visual media files.

[0140] The M candidate visual media files are obtained by matching the visual semantic vectors of the visual media files with the sentence semantic vectors of the search statement. This means that when recalling the M candidate visual media files, the degree of match between the visual semantics of the visual media files and the user's entire search statement is considered, without considering the degree of match between the visual semantics of the visual media files and the semantic subject of the search statement that the user visually focused on, resulting in the recall of some inaccurate photos and videos. Therefore, in this solution, the results of the first branch recall are fine-tuned by using the degree of match between the visual semantics of the visual media and the "semantic subject related to the visual content" to filter out inaccurate visual media.

[0141] Specifically, the semantic entities related to the visual content in the search statement and / or the search statement itself can be used as different dimensions to be matched. Different dimensions to be matched play different roles in fine-tuning, that is, the weights of different dimensions to be matched are different. The weight of each dimension to be matched is determined based on the degree of variation in the matching degree of the visual content of the M candidate visual media files on the dimension to be matched. The greater the degree of variation, the greater the role played by the dimension to be matched during fine-tuning, and therefore, the greater its weight. In this way, some inaccurate visual media such as photos and videos can be effectively excluded.

[0142] In one possible implementation, the method further includes:

[0143] Matching a search term in the search statement with a plurality of preset tags to determine whether the search term belongs to a semantic subject associated with the tag; the tag is used to describe visual content;

[0144] The semantic subject related to the tag in the search sentence is determined as the semantic subject related to the visual content in the search sentence.

[0145] In other words, a search term that matches one of the preset tags belongs to the semantic subject associated with the tag. Since tags are related to visual content, the semantic subject associated with the tag also belongs to the semantic subject associated with the visual content. In this solution, by presetting multiple tags, it is relatively easy to determine the semantic subject that the user is most visually interested in from the search statement.

[0146] In one possible implementation, the method further includes:

[0147] obtaining properties of the plurality of visual media files;

[0148] The plurality of second visual media files having attributes matching the semantic subject in the search statement that is unrelated to visual content are determined from the plurality of visual media files.

[0149] That is, recall is performed through the attributes of the visual media file (eg, time, location).

[0150] In one possible implementation, determining search results for the search statement based on the plurality of first visual media files includes:

[0151] If the search statement includes a time-related semantic subject, filtering the plurality of first visual media files according to the time-related semantic subject and the acquisition time attributes of the plurality of first visual media files;

[0152] If the search statement includes a semantic subject related to a location, filtering the plurality of first visual media files according to the semantic subject related to the location and the collection location attributes of the plurality of first visual media files;

[0153] If the search statement includes a semantic subject related to character relationships, filtering the plurality of first visual media files according to the semantic subject related to character relationships and character relationship attributes of the plurality of first visual media files; and / or

[0154] If the search statement includes a semantic subject related to a person's name, the plurality of first visual media files are filtered according to the semantic subject related to the person's name and the person's name attributes of the plurality of first visual media files.

[0155] In this solution, the recalled visual media are filtered by time, location, character relationship, and name.

[0156] In a fourth aspect, the present application provides an electronic device, comprising: a memory, a processor, and a display, wherein:

[0157] The memory is used to store programs;

[0158] The display is used to display the search page;

[0159] The processor is coupled to the memory and the display, and is configured to execute the program stored in the memory to implement any one of the methods described.

[0160] In a fifth aspect, the present application provides a computer-readable storage medium storing a computer program, wherein the computer program can implement any of the methods described above when executed by a computer.

[0161] In a sixth aspect, the present application provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, any of the methods described above is performed. BRIEF DESCRIPTION OF THE DRAWINGS

[0162] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0163] FIG1A is a set of interface diagrams of a search interface of a mobile phone entering a gallery application according to an embodiment of the present application;

[0164] FIG1B is a schematic diagram of a search interface after the search history is cleared according to an embodiment of the present application;

[0165] FIG1C is a set of interface diagrams related to searching in a gallery application according to an embodiment of the present application;

[0166] FIG1D is a schematic diagram of a search failure interface provided by an embodiment of the present application;

[0167] FIG2A is a schematic structural diagram of an electronic device provided in another embodiment of the present application;

[0168] FIG2B is a software structure block diagram of an electronic device provided in yet another embodiment of the present application;

[0169] FIG3A is another set of interface diagrams related to searching in a gallery application provided by an embodiment of the present application;

[0170] FIG3B is a set of interface diagrams related to searching in the negative one screen according to an embodiment of the present application;

[0171] FIG3C is a schematic diagram of a search result interface provided by an embodiment of the present application;

[0172] FIG3D is a second schematic diagram of a search result interface provided by an embodiment of the present application;

[0173] FIG3E is a third schematic diagram of a search result interface provided by an embodiment of the present application;

[0174] FIG3F is a fourth schematic diagram of a search result interface provided by an embodiment of the present application;

[0175] FIG4 is an interaction diagram of a visual media search method provided by an embodiment of the present application;

[0176] FIG5 is a flow chart of a visual media search method according to an embodiment of the present application;

[0177] FIG6 is a fifth schematic diagram of a search result interface provided in an embodiment of the present application. DETAILED DESCRIPTION

[0178] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships can exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0179] In the following, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the technical features indicated. Therefore, a feature specified as "first," "second," or "third" may explicitly or implicitly include one or more of the features.

[0180] First, the terms used in the embodiments of the present application are explained. It should be understood that this explanation is for a clearer understanding of the embodiments of the present application and does not necessarily constitute a limitation on the embodiments of the present application.

[0181] Visual media: refers to pictures or videos.

[0182] Semantic Subjects: Named Entity Recognition (NER) technology can identify entities with specific meanings, such as personal names (PERs) and place names (LOCs). In this solution, the identified entities with specific meanings are called semantic subjects.

[0183] Visual content-related and visual content-independent: Visual content refers to the objects displayed by visual media and the relationships between them. Computer vision can enable computers to have capabilities similar to human vision, including the perception, understanding, analysis, and interpretation of visual content. Currently, the Generative Pre-trained Transformer 4 (GPT-4) model can support inputting an image into the model and outputting human-language natural language describing the important information in the image.

[0184] In the context of image search in this solution, data related to visual media files that requires a model's natural image understanding capabilities to retrieve is referred to as "visual content-related." In other words, visual content is data that requires a natural image understanding model to retrieve. This solution also refers to data related to visual media files that can be retrieved without the model's image understanding capabilities as "visual content-independent." For example, when a terminal device captures a visual media file, it can retrieve and save the shooting location, shooting time, name, and file attributes.

[0185] For example, in "Photos taken in Beijing this year", "this year" (shooting time), "Beijing" (shooting location), and "photos" (file attributes) are all data that can be obtained and saved by the terminal device when collecting visual media files. Therefore, "this year", "Beijing", and "photos" are not related to visual content; in "The sky taken in Beijing this year", "sky" needs to be understood by the model's image understanding ability of the image or image frame of the visual media file, so "sky" is related to visual content.

[0186] Text semantic vector: A vector that can represent the semantic features of an entire sentence can be obtained by feeding the text into a text encoder. The text encoder can use models such as the Transformer commonly used in Natural Language Processing (NLP), but this solution does not impose any restrictions on this. In this solution, the text semantic vector obtained for a sentence is called a sentence semantic vector, the text semantic vector obtained for the semantic subject in the sentence is called a subject semantic vector, and the text semantic vector obtained for a label is called a label semantic vector.

[0187] Visual semantic vectors can be obtained by feeding images or frames from visual media files into an image encoder. This can be done using a common CNN (Convolutional Neural Network) model or a Vision Transformer (VIT) model, though this solution does not impose any restrictions.

[0188] Density-based clustering algorithm: It is based on a set of neighborhoods to describe the compactness of the sample set. (The first parameter ε, the second parameter MinPts) is used to describe the compactness of the sample distribution in the neighborhood. The first parameter ε is used to describe the neighborhood radius of the data point; the second parameter MinPts is used to describe the minimum number of data points in the neighborhood of the data point. Its representative algorithms include: DBSCAN (Density-Based Spatial Clustering of Application with Noise); DBSCAN algorithm is a relatively representative density-based clustering algorithm that can divide areas with sufficiently high density into clusters and can find clusters of any shape in a noisy spatial database;

[0189] Vector similarity: used to describe the degree of similarity between two vectors (e.g., between a sentence semantic vector and a visual semantic vector). In embodiments of the present application, the visual media whose visual content matches the search statement can be determined by comparing the similarity between the sentence semantic vector and the visual semantic vector. Generally, vector similarity can be calculated using the cosine similarity formula, but it can also be calculated using other methods.

[0190] In the prior art, mobile phones use gallery applications to manage users' visual media files such as pictures and videos (hereinafter referred to as "visual media"). Taking a photo taken with a mobile phone as an example, after the photo is taken, the gallery application can obtain and save attributes unrelated to the visual content, such as the photo's shooting location, shooting time, and photo title.

[0191] In actual applications, you can also input photos into the natural image understanding model when the phone is charging and the screen is off, so that the natural image understanding model can generate and save labels for the photos. The labels can be regarded as attributes of the photos related to the visual content. The labels can be: "sky", "cat", "dog", and so on. The gallery application can index the photos based on the attributes of the photos. After the index is established, the gallery application can provide corresponding search services to users. Specifically, users can search for pictures or videos by entering keywords in the gallery application. For example, users can enter keywords such as "Beijing", "sky", "National Day" in the search box provided by the gallery application. The gallery application matches the keywords entered by the user with the index of visual media such as pictures and videos in the gallery application to obtain search results.

[0192] The following describes the interfaces involved in the search process of the gallery application in the prior art with reference to the accompanying drawings:

[0193] As shown in (a) of FIG1A , the mobile phone can display a main interface 101, which can also be referred to as a desktop. The main interface 101 may include an icon 102 of a gallery application. The mobile phone may receive an operation in which a user clicks on the icon 102. In response to the operation, the mobile phone may start the gallery application and display an interface 103 as shown in (b) of FIG1A , wherein the interface 103 may be an album interface. It should be noted that in response to the user clicking on the icon 102, the mobile phone may start the gallery application and display the gallery's photo interface, which includes thumbnails of photos (i.e., pictures) in the gallery or a large picture of a particular photo. In the photo interface, in response to the user operating the "album" control, the above-mentioned album interface 103 is displayed.

[0194] As shown in (b) in Figure 1A, interface 103 includes multiple albums, among which the "All Photos" album includes 2023 photos, the "Camera" album includes 1502 photos and videos, the "Screenshots and Recordings" album includes 102 photos and videos, the "My Favorites" album includes 48 photos and videos, the "One Record, Multiple Gains" album includes 34 photos and videos, the "Video Editing" album includes 65 videos, the "Self-Created Album" includes 57 photos and videos, and the "Shared Album" includes 100 photos and videos.

[0195] As shown in (b) of Figure 1A , interface 103 may include a search box 104. The mobile phone may receive a user click on search box 104. In response, the mobile phone may display interface 105, shown in (c) of Figure 1A , which may be referred to as a search interface. Interface 105 may display photo classification information to the user. For example, in interface 105, the mobile phone categorizes photos by time, people, and objects. For example, within the time dimension, the mobile phone categorizes photos into three time periods: "This Month," "Last Month," and "This Year." The "This Month" album includes photos or videos taken this month, the "Last Month" album includes photos or videos taken last month, and the "This Year" album includes photos or videos taken this year. Within the person dimension, the mobile phone categorizes photos by different people, such as the four different people shown in interface 105. Within the object dimension, the mobile phone categorizes photos into "Landscape," "Animals," "Documents," and "Architecture." It should be noted that the above-mentioned classification dimensions may also be other, and no specific limitation is made here. In the interface 105, the user can also see these classification information without entering keywords.

[0196] Optionally, the interface 105 may further include a search history 107 and a "clear" 108 option. The search history includes keywords that the user has entered, such as "flowers", "coffee", "cat", etc. The mobile phone can receive an operation in which the user clicks on "clear" 108, and in response to this operation, the mobile phone can clear the search history. After the mobile phone clears the search history, the keywords that the user has entered are no longer displayed on the search interface 105. For example, in response to the user clicking on "clear" 108, as shown in FIG1B , the search history 107 and the "clear" 108 option are no longer displayed on the search interface 105, and the content displayed below moves up.

[0197] In response to the user inputting the keyword "sky" in interface 105, the mobile phone displays interface 109 as shown in (a) in Figure 1C. As shown in (a) in Figure 1C, there are 100 photos related to "sky" and 32 photos related to photos containing the word "sky". Among them, the 100 photos related to "sky" can be recalled because the tags of these 100 photos match "sky"; the 32 photos related to photos containing the word "sky" can be recalled because the OCR (Optical Character Recognition) recognition technology recognizes that these 32 photos contain the word "sky". In actual applications, the mobile phone can also associate the keywords entered by the user to obtain associated words, and search based on the associated words.

[0198] Interface 109 also displays some search results for the keyword "sky" and a "more" option 110 corresponding to the search results for the keyword "sky". The mobile phone receives a click operation by the user on the "more" option 110 and displays an interface 111 as shown in (b) of Figure 1C. Among them, interface 111 is used to display photos and videos in the search results for the keyword "sky". Optionally, photos and videos can be classified and displayed according to time. In addition, interface 111 also includes a return key 112 and a title 113. In response to the user's operation on the return key 112, the mobile phone can redisplay interface 109. Title 113 may include the keyword "sky".

[0199] That is, in existing gallery applications, users can enter simple keywords in the search box, such as "sky," "Beijing," "National Day," etc., and obtain corresponding search results. However, because the number of attributes of a photo is simple and limited, the mobile phone's ability to understand and associate search statements is also limited. If the user enters a more complex search statement in the search box, if the keywords in the search statement cannot match the attributes of the image or the text in the image, no photos will be found. In other words, existing gallery applications do not support search functions based on complex search statements. As shown in Figure 1D, when a user enters the more complex search statement "cooking tea around the fire" in the search box of interface 114, the mobile phone cannot understand the associated words of "cooking tea around the fire." Since the photo does not have attributes that match "cooking tea around the fire" or its associated words, the search result is displayed as "no pictures."

[0200] However, in actual applications, users have a strong demand for the function of searching for images based on complex search statements. This is because users can use complex search statements to more comprehensively describe the images or videos they want, thereby achieving accurate searches. In order to meet this user demand, an embodiment of the present application provides a visual media search method. The method can be applied to electronic devices, which can be mobile phones, tablet computers, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), and other terminal devices.

[0201] 2A shows a schematic diagram of the structure of an electronic device 200. The electronic device 100 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, an earphone interface 270D, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display 294, and a subscriber identification module (SIM) card interface 295.

[0202] The sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, an air pressure sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.

[0203] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0204] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.

[0205] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.

[0206] Processor 210 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 210 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 210. If processor 210 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 210 latency, and thus improves system efficiency.

[0207] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0208] It is understood that the interface connection relationship between the modules illustrated in the embodiment of the present application is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may also adopt a different interface connection method from the above embodiment, or a combination of multiple interface connection methods.

[0209] Electronic device 200 implements display functionality through a GPU (Graphics Processing Unit), display screen 294, and an application processor. The GPU connects display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 210 may include one or more GPUs that execute program instructions to generate or modify display information.

[0210] Display screen 294 is used to display images, videos, and the like. Display screen 294 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, electronic device 100 can include one or N display screens 294, where N is a positive integer greater than one.

[0211] The electronic device 200 can implement a shooting function through an ISP, a camera 293, a video codec, a GPU, a display screen 294, and an application processor.

[0212] The camera 293 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 100 may include 1 or N cameras 293, where N is a positive integer greater than 1.

[0213] Video codecs are used to compress or decompress digital video. Electronic device 200 may support one or more video codecs. This allows electronic device 200 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.

[0214] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU enables intelligent cognitive applications in electronic device 200, such as image recognition, face recognition, speech recognition, and text comprehension.

[0215] The external memory interface 220 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 200. The external memory card communicates with the processor 210 via the external memory interface 220 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.

[0216] The internal memory 221 can be used to store computer executable program codes, which include instructions. The internal memory 221 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device 200 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 221 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 210 executes various functional applications and data processing of the electronic device 200 by running instructions stored in the internal memory 221 and / or instructions stored in a memory provided in the processor.

[0217] FIG2B is a block diagram of the software structure of the electronic device 200 according to an embodiment of the present application. The software system of the electronic device 200 can adopt a layered architecture, which divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. Taking the Android system as an example, in some embodiments, the Android system can be divided into four layers, namely, the application layer, the application framework layer, the Android runtime (Android runtime) and the system library, and the kernel layer, from top to bottom.

[0218] As shown in FIG2B , the application layer may include applications such as a gallery service module, a search module, a multimodal understanding module, a natural language understanding module, and a camera application.

[0219] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0220] The system library can include multiple functional modules, such as the surface manager, media libraries, 3D graphics libraries (such as OpenGL ES), and 2D graphics engines (such as SGL). The media library supports playback and recording of various common audio and video formats, as well as static image files. The media library supports a variety of audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0221] The kernel layer is the layer between hardware and software.

[0222] The following describes the workflow of the software and hardware of the electronic device 200 in conjunction with capturing a photo scene.

[0223] When touch sensor 280K receives a touch operation, a corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including touch coordinates, touch operation timestamp, and other information). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. For example, if the touch operation is a single-touch operation and the control corresponding to the single-touch operation is a control of the camera application icon, the camera application calls the interface of the application framework layer to start the camera application, and then starts the camera driver by calling the kernel layer, and captures a still image or video through the camera 293.

[0224] Taking a mobile phone as an example, the visual media search method provided by the embodiment of the present application is introduced. The visual media search method provided by the embodiment of the present application can be applied in applications such as gallery applications and file management applications.

[0225] The interface and search logic involved in the visual media search method provided in the embodiment of the present application will be described below with reference to the accompanying drawings.

[0226] As shown in (a) of FIG3A , interface 301 displays a search history 303 and a "Clear" option 304. The search history includes search statements that the user has entered, such as "watching the sunrise on the top of a mountain" and "photographing the Great Wall in Beijing this year." Other content that can be displayed in interface 301 can refer to the relevant display content of interface 305 above and will not be repeated here. Interface 301 can be referred to as a search interface. The mobile phone can display interface 301 in response to a user clicking on the search box 104 in interface 103 shown in (b) of FIG1A .

[0227] As shown in interface 305 (i.e., the search results interface) in Figure 3A (b), a user enters the search phrase "cooking tea around the fire" in search box 306 of interface 305, and the phone searches for 239 images. Interface 305 displays a portion of the search results for the phrase "cooking tea around the fire" (e.g., thumbnails of eight images) and a "More" option 307 corresponding to the search results for the phrase "cooking tea around the fire." In response to the user's click on "More" option 307, the phone displays interface 308, as shown in Figure 3A (c). Interface 308 displays images from the search results for the phrase "cooking tea around the fire" in descending order, based on the degree of match between the image's visual content and the search phrase (specifically, the similarity between the image's visual semantic vector and the sentence semantic vector of the search phrase). The user can also swipe up on interface 308 to view undisplayed images.

[0228] Currently, mobile phones also have a negative first screen and a pull-down search interface. It is understandable that the negative first screen can be the leftmost split screen of the electronic device, used to provide users with functions such as search and quick services. Among them, the negative first screen can also be used to display notification messages that need to be pushed to the user, such as user-subscribed application messages, real-time hot search messages, segment selection, itinerary information, etc. The pull-down search interface is an interface displayed in response to the user's pull-down operation on the main interface. This interface is used to provide users with functions such as search and application suggestions. This interface can be the same interface as interface 315 in Figure 3B.

[0229] The following will be introduced by taking the negative one screen as an example. When the user needs to view the negative one screen of the mobile phone, the user can slide the screen of the mobile phone to make the electronic device display the negative one screen.

[0230] Exemplarily, as shown in (a) of FIG3B , the mobile phone may receive a first operation performed by the user on the interface 309 of the mobile phone (which may be called the desktop). Exemplarily, the first operation may be a right sliding operation as shown in (a) of FIG3B . In response to the first operation, the mobile phone may display a negative screen 310 as shown in (b) of FIG3B . Among them, the negative screen 310 may include: a search box 311, a quick service 312, a default card 313, a recommended card 314, and the like. The quick service 312 may be a quick entry to a page or function of an application, such as: scan, payment code, ride code, etc.; the default card may be: a gallery card, a remaining power card, etc.; the recommended card may be a recommended application card.

[0231] The mobile phone receives a click operation by the user on the search box 311 on the negative first screen 310, and displays an interface 315 as shown in (c) of Figure 3B. The interface 315 may include: a search box 316, application suggestions, and a search history 217. Application suggestions include: icons of each application recommended for use. The interface 215 may also include: a search history 217 and its corresponding "clear" option 318. In response to the user's triggering operation on the "clear" option 318, the search history 317 and the "clear" option 318 are no longer displayed on the interface 315. In addition, the search box 316 may display hot news headlines, such as: "Tianjin Marathon".

[0232] As shown in interface 319 (d) of FIG3B , the search box of interface 319 displays the user-entered search query "National Day photo taken of the Great Wall in Beijing." Interface 319 also displays a preview area 322 of the gallery's search results for the search query "National Day photo taken of the Great Wall in Beijing," and a corresponding "Search in App" option 323. In response to a user triggering operation on preview area 322, the user enters a photo details interface provided by the gallery application, allowing the user to flip through pages of search results for the search query "National Day photo taken of the Great Wall in Beijing." In response to a user triggering operation on "Search in App" option 323, the phone displays interface 324 provided by the gallery application, as shown in FIG3B (e). Interface 324 displays a portion of the search results for the search query "National Day photo taken of the Great Wall in Beijing," as well as a "More" option corresponding to the search results for the search query "National Day photo taken of the Great Wall in Beijing." In response to a user triggering operation on the "More" option, the phone displays a search results details interface, which displays images from the search results for the search query "National Day photo taken of the Great Wall in Beijing." An online search option 321 may also be displayed in the interface 319. In response to the user triggering the online search option 321, the mobile phone displays a search webpage and displays online search results on the search webpage.

[0233] In interface 325 shown in Figure 3C, the user enters the search sentence "The Great Wall photographed on National Day last year" in the search box of interface 325, and the mobile phone searches for 419 pictures. The mobile phone displays the thumbnails of each picture or video in the search results of the search sentence "The Great Wall photographed on National Day last year" on interface 325. Among them, the shooting time of the picture or video corresponding to thumbnail A is 23:22 on September 30, 2022; the shooting time of the picture or video corresponding to thumbnail B is 00:19 on October 1, 2022. Among them, the time search range corresponding to National Day is from 00:00 on October 1, 2022 (that is, the first time point) to 00:00 on October 8, 2022 (that is, the second time point). The shooting time of the picture or video corresponding to thumbnail A is before 00:00 on October 1, 2022. The shooting time of the picture or video corresponding to thumbnail B is 00:19 on October 1, 2022, which is between 00:00 on October 1, 2022 and 00:00 on October 8, 2022. Among them, the shooting time of the pictures or videos corresponding to thumbnail A and thumbnail B is relatively close.

[0234] As shown in interface 326 of Figure 3D, for the search statement "the sky taken on National Day last year", the shooting time of the picture or video corresponding to the displayed thumbnail C is 22:19 on October 7, 2022; the shooting time of the picture or video corresponding to the displayed picture D is 01:24 on October 8, 2022. Among them, the time search range corresponding to National Day is from 00:00 on October 1, 2022 (that is, the first time point) to 00:00 on October 8, 2022 (that is, the second time point). The shooting time of the picture or video corresponding to thumbnail D is after 00:00 on October 8, 2022. The shooting time of the picture or video corresponding to thumbnail C is between 00:00 on October 1, 2022 and 00:00 on October 8, 2022. Among them, the pictures or videos corresponding to thumbnail C and thumbnail D are shot at a relatively close time.

[0235] In practical applications, the search statement may include not only a time search range (such as National Day) but also a first keyword (such as the Great Wall or the sky). Then, for the search statement, the visual media file corresponding to the search result matches the first keyword.

[0236] When the first keyword belongs to a semantic subject that is unrelated to the visual content (eg, a location), the first keyword may be matched with attributes of the visual media to determine a visual media file that matches the first keyword.

[0237] When the first keyword belongs to a semantic subject related to the visual content (such as the Great Wall or the sky), the first keyword is matched with the tag of the visual media file to determine the visual media file whose visual content matches the first keyword; or, the text semantic vector of the first keyword is matched with the visual semantic vector of the visual media file to determine the visual media file whose visual content matches the first keyword (including the first, second, and third visual media files mentioned above).

[0238] The thumbnails are proportional thumbnails of corresponding visual media files or proportional thumbnails of any image frame.

[0239] In interface 327 shown in FIG3E , for the search phrase "sky photographed during last year's National Day," two thumbnails E and F are displayed, sorted in descending order by the degree of match between their visual content and the first keyword "sky." The visual content of the visual media file corresponding to thumbnail E has a match of 0.87 with "sky," while the visual content of the visual media file corresponding to thumbnail F has a match of 0.76 with "sky."

[0240] In interface 328 shown in FIG3F , for the search phrase "Photos taken in Tianjin's Nankai District during last year's National Day," two thumbnails J and H are displayed, sorted in descending order by the degree of matching their respective visual media file's location attributes with the first keyword, "Tianjin's Nankai District." The location attribute of the visual media file corresponding to thumbnail J matches "Tianjin's Nankai District" with a degree of 0.87, while the location attribute of the visual media file corresponding to thumbnail H matches "Tianjin's Nankai District" with a degree of 0.8.

[0241] Figure 4 is an interactive diagram of the visual media search method provided by an embodiment of the present application. As shown in Figure 4 , the mobile phone is provided with: a gallery service module (also known as a gallery application) 41 , a search module 42 , a multimodal understanding module 43 , and a natural language understanding module 44 .

[0242] As shown in FIG4 , the visual media search method provided in the embodiment of the present application can be divided into two stages: an index building stage and a search stage.

[0243] The index building phase includes the following steps:

[0244] S401. Add and / or modify visual media and its attributes.

[0245] In S401 above, for newly added visual media, the gallery application can automatically generate attributes unrelated to the visual content, including but not limited to: acquisition location, acquisition time, and visual media name. For captured videos or images, for example, the acquisition location refers to the location where the video was taken, and the acquisition time refers to the time the video was taken. For screenshots, for example, the acquisition location refers to the location where the screenshot was taken, and the acquisition time refers to the time the screenshot was taken. For downloaded videos or images, for example, the acquisition location refers to the download location, and the acquisition time refers to the download time.

[0246] Users can add new visual media by shooting, downloading, or taking screenshots. Additionally, users can modify existing visual media, including but not limited to beautification, custom naming, and adding watermarks.

[0247] S402: Store the visual media and its attributes.

[0248] In response to the above addition or modification operations, the gallery service module 41 can store the visual media and its attributes locally on the mobile phone. In actual applications, with the user's authorization, the mobile phone can store the locally stored visual media and its attributes to the cloud to reduce the local storage pressure of the mobile phone.

[0249] S403: Request visual semantic understanding of the visual media.

[0250] Since visual semantic understanding requires a large amount of computing resources, in order not to affect user usage, the above step S403 can be performed when the mobile phone is in charging and screen-off state.

[0251] The gallery service module 41 may request the multimodal understanding module 43 to perform visual semantic understanding on the newly added or modified visual media to obtain a visual semantic vector of the visual media.

[0252] The multimodal understanding module 43 may perform visual semantic understanding on the visual media based on the multimodal model to obtain a visual semantic vector of the visual media.

[0253] The multimodal model is used not only to: perform visual semantic understanding of visual media to obtain visual semantic vectors for the visual media; but also to: perform semantic understanding of search sentences to obtain sentence semantic vectors for the search sentences; perform semantic understanding of the rewritten search sentences in the subsequent text to obtain sentence semantic vectors for the rewritten search sentences; perform semantic understanding of the semantic subject in the search sentence to obtain the subject semantic vector of the semantic subject; and perform semantic understanding of the labels in the visual media to obtain the label semantic vectors of the labels. The multimodal model can be trained based on training samples.

[0254] Exemplarily, the multimodal model can be CLIP (Contrastive Language-Image Pre-training, a pre-training model based on contrasting text-image pairs). The CLIP model can be used to map visual media and text (i.e., search statements, rewritten search statements, semantic entities, and labels) into a unified vector space to understand the relationship between different modal resources in text and visual terms, which can then be used for image retrieval. That is, in embodiments of the present application, the CLIP model can be used to match visual media files with text.

[0255] The multimodal model maps visual media and text into vectors of the same dimensionality. This means that the dimensionality of the visual semantic vector of the visual media is the same as the dimensionality of the text semantic vector (for example, the sentence semantic vector of a search query). The multimodal model includes the image encoder and text encoder described above.

[0256] S404: Return the visual semantic vector of the visual media.

[0257] The multimodal understanding module 44 returns the visual semantic vector of the visual media to the gallery service module 41 .

[0258] S405: Store the visual semantic vector of the visual media.

[0259] The gallery service module 41 may store the visual semantic vectors of the visual media locally.

[0260] S406: Send attribute information of the visual media and its visual semantic vector.

[0261] Exemplarily, the gallery service module 41 may store the visual semantic vectors of the visual media returned by the multimodal understanding module 44 , and then send the attributes of the visual media and their visual semantic vectors in batches to the search module 42 so that the search module 42 can construct an index of the visual media.

[0262] S407, build index

[0263] The index of visual media constructed by the search module 42 may include: attributes of the visual media and visual semantic vectors of the visual media.

[0264] The search phase includes the following steps:

[0265] S408: Input a search statement.

[0266] The user can enter a search statement through the search interface provided by the gallery service module 41, such as the interface 301 shown in (a) of FIG3A . For example, as shown in (b) of FIG3A , the user enters “cooking tea around the stove” in the search box 306.

[0267] S409: Send the search statement.

[0268] After receiving the search statement input by the user, the gallery service module 41 sends the search statement to the search module 42 for searching.

[0269] S410: Requesting semantic subject recognition for the search statement.

[0270] Search module 42 requests natural language understanding module 44 to perform semantic subject recognition on the search statement to obtain the semantic subject contained in the search statement. Natural language understanding module 44 performs semantic subject recognition based on a natural language understanding model. Specifically, named entity recognition (NER) technology can be used to perform semantic subject recognition on the search statement to obtain the semantic subject contained in the search statement. In the embodiment of the present application, the semantic subject can also be referred to as an entity.

[0271] Named entity recognition technology can be used to identify semantic subjects related to time, location, and tags in search statements. Among them, semantic subjects related to time and location are semantic subjects that are unrelated to visual content; the tags of visual media files are data that need to be obtained through natural image understanding of the model. Therefore, semantic subjects related to tags are semantic subjects related to visual content. In practical applications, based on actual experience, multiple tags that users are more concerned about can be counted, such as: "sky", "cat", "dog", "birthday", "child", etc. These tags are used to describe visual content. Subsequently, named entity recognition technology can match the keywords (or search terms) in the search statement with multiple pre-set tags to determine whether the keyword belongs to a semantic subject related to the tag.

[0272] For example, named entity recognition technology is used to perform semantic subject recognition on the search sentence "the sky photographed in Beijing during National Day", and it is determined that "National Day" belongs to the semantic subject related to time, "Beijing" belongs to the semantic subject related to location, and "sky" belongs to the semantic subject related to label.

[0273] S411. Return the semantic subject.

[0274] The natural language understanding module 44 returns the recognized semantic subject to the search module 42 .

[0275] S412: Request semantic understanding of the search statement.

[0276] The search module 42 can send the search statement to the multimodal understanding module 43, which performs semantic understanding on the search statement to obtain a sentence semantic vector (i.e., a first sentence semantic vector) of the search statement. The specific semantic understanding process can be found in the corresponding content of the above embodiments and will not be repeated here.

[0277] It should be noted that when the search statement includes a semantic subject related to the visual content (i.e., a semantic subject related to the tag), the search module 42 may also send the semantic subject related to the visual content to the multimodal understanding module 43, so that the multimodal understanding module 43 can perform semantic understanding on the semantic subject related to the visual content to obtain a subject semantic vector for the semantic subject. Continuing with the above example, "sky" belongs to the semantic subject related to the visual content, and the multimodal understanding module 43 can perform semantic understanding on "sky" to obtain a subject semantic vector corresponding to "sky".

[0278] S413. Return the text vector.

[0279] The multimodal understanding module 43 may return the sentence semantic vector of the search sentence to the search module 42 .

[0280] S414: Perform recall based on the attributes of the visual media and the visual semantic vector of the visual media.

[0281] The search module 42 includes different branch search methods:

[0282] Exemplarily, the first branch involves recalling visual media based on their visual semantic vectors. Specifically, the visual semantic vectors of each visual media from multiple visual media stored on the mobile phone are obtained; the vector similarity between the sentence semantic vector of the search statement and the visual semantic vectors of each visual media is calculated; based on the vector similarity, M candidate visual media (i.e., M candidate visual media files) whose visual semantic vectors match the sentence semantic vector are determined from the multiple visual media stored on the mobile phone; where M is an integer greater than or equal to 1; these M visual media can be used as the visual media set recalled by the first branch.

[0283] Exemplarily, the visual media whose vector similarity is greater than a preset similarity threshold is used as the visual media whose visual semantic vector matches the sentence semantic vector.

[0284] Exemplarily, multiple visual media are sorted from high to low according to vector similarity, and the top F (F≥1) visual media are used as the visual media whose visual semantic vectors match the semantic vector of the sentence.

[0285] Exemplarily, Q (Q ≥ 1) visual media whose vector similarity is greater than a preset similarity threshold are determined from multiple visual media stored via a mobile phone; if Q is greater than or equal to a preset number threshold D, these Q visual media can be sorted from high to low according to vector similarity; the top D visual media are used as visual media whose visual semantic vectors match the semantic vector of the sentence; if Q is less than the preset number threshold D, these Q visual media can be directly used as visual media whose visual semantic vectors match the semantic vector of the sentence.

[0286] It should be noted that the multiple visual media stored on the mobile phone may include: visual media stored locally on the mobile phone and / or visual media stored by the mobile phone in the cloud (e.g., cloud storage space applied for by the mobile phone). To protect user privacy, the multiple visual media stored on the mobile phone are all stored on the mobile phone.

[0287] It should be noted that the first branch of the recall process does not understand the attributes of the visual media, but only understands the overall visual semantic information of the visual media.

[0288] Exemplarily, the second branch is: recall based on attributes of the visual media.

[0289] Specifically, the method obtains the attributes of each visual medium from a plurality of visual media stored on the mobile phone; matches the semantic subject in the search statement with the attributes of each visual medium to determine the visual medium matched by the semantic subject (i.e., the second visual medium); and uses the visual media matched by the semantic subject as the visual media set recalled by the second branch. When the search statement contains only one semantic subject, the visual media set recalled by the second branch includes the visual media matched by the single semantic subject; when the search statement contains multiple semantic subjects, the visual media set recalled by the second branch includes the visual media matched by each of the multiple semantic subjects.

[0290] For example, for a semantic subject related to time, the time range corresponding to the semantic subject can be determined (i.e., the time search range); this time range can be matched with the acquisition time attribute of each visual media to determine the visual media whose acquisition time falls within this time range; and the visual media whose acquisition time falls within this time range are used as the visual media matched by the semantic subject. For example, if the semantic subject is "National Day" and its corresponding time range is "October 1 to October 7", the acquisition time of picture 1 is "October 2", and the acquisition time of picture 2 is "October 8", then according to the above matching method, picture 1 matches the semantic subject "National Day", while picture 2 does not match the semantic subject "National Day".

[0291] For example, for a semantic subject related to a location, the geographic scope (i.e., geographic search scope) corresponding to the semantic subject can be determined; the geographic scope can be matched with the collection location attribute of each visual media to determine the visual media whose collection location is within the geographic scope; and the visual media whose collection location is within the geographic scope can be used as the visual media matched by the semantic subject. For example, if the semantic subject is "Beijing" and its corresponding geographic scope is the entire Beijing city, the collection location of picture 3 is "Xicheng District, Beijing", and the collection location of picture 4 is "Nankai District, Tianjin", then according to the above matching method, picture 3 matches the semantic subject "Beijing", while picture 4 does not match the semantic subject "Beijing".

[0292] Exemplarily, for a semantic subject associated with a label, the subject semantic vector of the semantic subject and the label semantic vector of each visual media label can be obtained; the vector similarity between the subject semantic vector of the semantic subject and the label semantic vector of each visual media label can be calculated; and based on the vector similarity, the visual media that matches the semantic subject can be determined. For example, if the semantic subject is "human cub" and picture 5 has a label "child", through calculation, it is found that the subject semantic vector of "human cub" is similar to the label semantic vector of "child", that is, picture 5 matches the semantic subject "human cub". In actual applications, after obtaining the visual media set recalled by the first branch and the visual media set recalled by the second branch, a candidate visual media set can be determined based on the visual media set recalled by the first branch and the visual media set recalled by the second branch. In an optional embodiment, the union or intersection of the visual media set recalled by the first branch and the visual media set recalled by the second branch can be used as the candidate visual media set.

[0293] In actual applications, when users search for images on their phones, they sometimes focus on the visual semantics of the image, sometimes on attributes such as the location and time of the photo, or sometimes on both. For example, when a user searches for "photos taken today," the user focuses on the time the image was taken. When a user searches for "sky taken today," the user focuses not only on the time the image was taken but also on the visual semantics of the image, that is, whether the image is of the sky. When a user searches for "photos taken while walking in Beijing," the user focuses on the location of the image. When a user searches for "photos of a walk in Beijing," the user focuses not only on the location of the image but also on the visual semantics of the image, that is, whether the image is of a walk.

[0294] Taking the two search statements "photos taken in Beijing this year" and "the sky taken in Beijing this year" as examples, referring to the above description, the semantic proportion of the relevant visual content in the search statement "the sky taken in Beijing this year" is greater than the semantic proportion of the relevant visual content in the search statement "photos taken in Beijing this year". Obviously, for the search statement "photos taken in Beijing this year", it is more appropriate to use the visual media set recalled by the second branch as the candidate visual media set. In this way, the intersection of the visual media matching "this year" and the visual media matching "Beijing" in the candidate visual media set can be used as the final search result. The collection location of each visual media in the final search result is Beijing and the collection time is this year, which meets the user's search requirements. If the visual media set recalled by the first branch is used as the candidate visual media set, the following situation may occur: there is a picture of "a girl holding a camera and taking a picture" in the visual media set recalled by the first branch, but the photo was taken in Shanghai and was taken last year. Since the search phrase "photos taken in Beijing this year" contains the word "photographing," and the image "girl holding up a camera and taking photos" contains the action of "photographing," there is a certain degree of similarity between the sentence semantic vector of the search phrase "photos taken in Beijing this year" and the visual semantic vector of the image "girl holding up a camera and taking photos." This means that the image "girl holding up a camera and taking photos" is likely to be recalled. In other words, when the semantic content of the visual content accounts for a low proportion, it is inappropriate to use the visual media set recalled by the first branch as a candidate visual media set.

[0295] In order to solve the above problem, in an optional embodiment, the proportion of semantics related to visual content in the search statement can be determined; based on the proportion of semantics related to visual content in the search statement, the first extraction ratio corresponding to the visual media set recalled by the first branch and the second extraction ratio corresponding to the visual media set recalled by the second branch are determined; when the semantic proportion is greater than or equal to a preset proportion threshold, the first extraction ratio is greater than the second extraction ratio. For example, if the semantic proportion is S, the first extraction ratio is: β*S, and the second extraction ratio is: 1-β*S, where the value of β can be set according to actual needs, and the embodiment of the present application does not specifically limit this.

[0296] In another optional embodiment, the proportion of semantics related to visual content in the search statement can be determined; when the proportion of semantics related to visual content in the search statement is greater than or equal to a preset proportion threshold, the visual media set recalled by the first branch is used as a candidate visual media set; when the proportion of semantics related to visual content in the search statement is less than the preset proportion threshold, the visual media set recalled by the second branch is used as a candidate visual media set. When the proportion of semantics related to visual content in the search statement is greater than or equal to the preset proportion threshold, it means that the search statement is a high-semantic search statement; when the proportion of semantics related to visual content in the search statement is less than the preset proportion threshold, it means that the search statement is a low-semantic search statement. The specific process will be described in detail in the following embodiment in conjunction with Figure 5.

[0297] S415, visual media filtering.

[0298] The search module 42 may perform semantic subject filtering, spatiotemporal filtering, and / or character relationship filtering on the candidate visual media set. The specific filtering method will be described in detail in the following embodiment in conjunction with FIG. 5 .

[0299] S416, visual media sorting.

[0300] Search module 42 ranks the candidate visual media sets.

[0301] The specific sorting method will also be described in detail in the following embodiments.

[0302] S417. Return search results.

[0303] The search module 42 sends the sorted search results to the gallery service module 41 .

[0304] Specifically, the search module 42 may select the top P (P≧1) visual media and their ranking information as the final search results.

[0305] S418. Display search results.

[0306] The gallery service module 41 may display the search results to the user, for example, by displaying the search results through the interface 305 in FIG. 3A or the interface 308 in FIG. 3A .

[0307] As shown in interface 305 of Figure 3A , the search results include 239 images, but interface 305 only displays thumbnails of eight of the images in the search results. To view all 239 images, the user can click "More" in interface 305 . In response, the phone displays interface 308 of Figure 3A . In one example, the order in which images are displayed in the search results is related to their matching degree. For example, images displayed earlier in the search results have a matching degree greater than or equal to that of images displayed later in the search results. The calculation of matching degrees and the sorting method will be described in detail in the following embodiments.

[0308] The search process performed by the search module 42 of the present application will be described in detail below with reference to FIG5 :

[0309] 501. Receive a search statement.

[0310] 502. Execute a search process corresponding to the search statement based on the visual semantic vector of the visual media.

[0311] Step 502 is executed to obtain the visual media set recalled by the first branch. For details, please refer to the search process corresponding to the first branch described above.

[0312] For example, let us assume that the multimodal model can encode the visual media and the search sentence into a k-dimensional vector space. The search sentence is Q, and the corresponding sentence semantic vector V is Q ={α z}, z=1,2,…,Z, a total of M pictures are recalled, and their vector representation is:

[0313] Among them, the i-th row corresponds to the visual semantic vector of the i-th image, and the value range of i is [1, Mm].

[0314] 503. Semantic subject identification.

[0315] The specific process of semantic subject recognition for the search statement can be found in the corresponding content of the above embodiment and will not be repeated here.

[0316] 504. Perform a search based on properties of the visual media.

[0317] Specifically, the semantic subjects in the search statement that are not related to the time content are matched with the attributes of the visual media to obtain the visual media set recalled by the second branch. For details, please refer to the search process corresponding to the second branch above.

[0318] 505. Rewrite the search statement.

[0319] Specifically, the semantic entities unrelated to visual content in the search statement are deleted to obtain the rewritten search statement. The semantic entities unrelated to visual content specifically refer to: semantic entities related to time and semantic entities related to location.

[0320] Exemplarily: for the search statement "the sky photographed this year", where "this year" is a semantic entity related to time, its rewritten search statement is "the photographed sky".

[0321] In practical applications, after deleting the semantic entities unrelated to visual content, there may be some redundant stop words. For example: for the search statement "the sky photographed in Beijing this year", where "this year" is a semantic entity related to time and "Beijing" is a semantic entity related to location, after deleting "this year" and "Beijing", the stop word "in" becomes a redundant word, so it also needs to be deleted. Specifically, the semantic entities unrelated to visual content and their related stop words in the search statement are deleted to obtain the rewritten search statement. Exemplarily, the rewritten search statement corresponding to the search statement "the sky photographed in Beijing this year" is "the photographed sky".

[0322] It should be noted that there is no order restriction in the execution of the above steps 502, 503, and 505. In an optional example, to improve efficiency, these three steps can be executed simultaneously.

[0323] 506. Perform the search process corresponding to the rewritten search statement based on the visual semantic vectors of the visual media.

[0324] Performing step 506 obtains the set of visual media recalled by the third branch.

[0325] Specifically, obtain the visual semantic vectors of each visual media among the multiple visual media stored in the mobile phone; calculate the vector similarity between the sentence semantic vector of the rewritten search statement (i.e., the second sentence semantic vector) and the visual semantic vectors of each visual media; according to the vector similarity, determine multiple visual media (i.e., reference visual media) whose visual semantic vectors match the sentence semantic vector from the multiple visual media stored in the mobile phone; these multiple visual media can be used as the set of visual media recalled by the third branch.

[0326] Among them, the specific implementation process of the step "according to the vector similarity, determine multiple visual media whose visual semantic vectors match the sentence semantic vector from the multiple visual media stored in the mobile phone" can refer to the corresponding content in the above embodiments and will not be elaborated here.

[0327] Exemplarily, the rewritten search statement is Q′, and the corresponding vector V Q′ ={α′ z}, z = 1, 2,..., Z, and a total of M′ picture sets R′ are recalled, and its vector representation is:

[0328] Among them, the i-th row corresponds to the visual semantic vector of the i-th image, and the value range of i is [1, M′].

[0329] 507. Calculate the proportion of semantics related to visual content in search sentences.

[0330] The following describes a method for determining the semantic proportion related to visual content:

[0331] 5071. Determine the representative visual semantic vector V corresponding to the visual media set recalled by the first branch I (i.e. the first representative visual semantic vector) and the representative visual semantic vector V corresponding to the visual media set recalled by the third branch I′ (That is, the second represents the visual semantic vector).

[0332] 5072. Determine the sentence semantic vector V Q and the visual semantic vector V I The difference between them is used to obtain the difference vector (V Q -V I )(that is, the first difference vector).

[0333] 5073. Determine the sentence semantic vector V Q′ and the visual semantic vector V I′ The difference between them is used to obtain the difference vector (V Q′ -V I′ )(that is, the second difference vector).

[0334] 5074. According to the difference vector (V Q -V I ) and the difference vector (V Q′ -V I′ ) to determine the semantic proportion related to visual content in the search sentence.

[0335] Among them, the proportion of semantics related to visual content is positively correlated with the similarity of the vector.

[0336] In the above 5071, in one example, the vector similarity between the visual semantic vector of each visual media in the visual media set recalled by the first branch and the sentence semantic vector of the search sentence can be obtained; the visual media in the visual media set recalled by the first branch are sorted from high to low according to the vector similarity; the average vector of the visual semantic vectors of the top T (T ≥ 1) visual media is used as the representative visual semantic vector V I .

[0337] Continuing with the above example, it represents the visual semantic vector V I for:

[0338] The vector similarity between the visual semantic vector of each visual media in the visual media set recalled by the third branch and the sentence semantic vector of the rewritten search sentence can be obtained; the visual media in the visual media set recalled by the third branch are sorted from high to low according to the vector similarity; the average vector of the visual semantic vectors of the top H (H ≥ 1) visual media is taken as the representative visual semantic vector V I′ .

[0339] Continuing with the previous example: represents the visual semantic vector V I′ for:

[0340] The values ​​of H and T may be the same or different, and this embodiment of the present application does not impose any specific limitation thereto. In another example, a clustering algorithm may be used to cluster the visual media in the visual media set recalled by the first branch based on the vector similarity between the visual semantic vector of each visual media in the visual media set recalled by the first branch and the sentence semantic vector of the search sentence, thereby obtaining cluster centers; and the visual semantic vector of the cluster center is used as the first representative visual semantic vector.

[0341] According to the vector similarity between the visual semantic vector of each visual media in the visual media set recalled by the third branch and the semantic vector of the second sentence of the rewritten search statement, the visual media in the visual media set recalled by the third branch can be clustered using a clustering algorithm to obtain the cluster center point; the visual semantic vector of the cluster center point is used as the second representative visual semantic vector.

[0342] In the embodiment of the present application, the visual semantic vector is used to represent the visual semantics of the entire set. The specific clustering algorithm can be selected according to actual needs, and the embodiment of the present application does not impose any limitation on this.

[0343] In the above 5074, in an optional embodiment, the difference vector (V Q -V I ) and the difference vector (V Q′ -V I′ ) is directly used as the semantic proportion related to visual content in the search sentence.

[0344] Among them, the difference vector (V Q -V I ) and the difference vector (V Q′ -V I′ ), the greater the vector similarity between them, the greater the semantic proportion of the visual content in the search sentence; conversely, the smaller the semantic proportion of the visual content in the search sentence.

[0345] It should be noted that in the embodiment of the present application, the method for calculating the proportion of semantics related to visual content in the search statement is based on the word analogy characteristics of the distributed representation vector (Distribution Representation), that is, the additivity of word meaning is directly reflected in the additivity of the distributed representation vector.

[0346] When the proportion of semantics related to visual content in the search statement is less than the preset proportion threshold, step 508 is subsequently executed.

[0347] When the proportion of semantics related to visual content in the search statement is greater than or equal to a preset proportion threshold, steps 509 and 510 are subsequently performed.

[0348] 508. Determine the visual media recalled by the second branch as a candidate set.

[0349] 509. The visual media recalled by the first branch are determined as a candidate set.

[0350] 510. Visual media filtering.

[0351] In order to improve the search accuracy, one or more of the following processes can be performed on the visual media recalled by the first branch: semantic subject filtering, spatiotemporal filtering, name filtering, and character relationship filtering. Among them, semantic subject filtering refers to filtering the visual media recalled by the first branch based on the degree of matching between the visual media recalled by the first branch and the semantic subject related to the visual content in the search statement. Spatiotemporal filtering includes: time filtering and space (i.e., location) filtering. Name filtering refers to filtering the visual media recalled by the first branch based on the degree of matching between the name attributes of the visual media recalled by the first branch and the semantic subject related to the name in the search statement. Character relationship filtering refers to filtering the visual media recalled by the first branch based on the degree of matching between the character relationship attributes of the visual media recalled by the first branch and the semantic subject related to the character relationship in the search statement. In actual applications, in addition to the above-mentioned attributes such as the collection location, collection time, and visual media name, the visual media stored in the mobile phone can also include name attributes and character relationship attributes manually entered by the user. Therefore, in practical applications, named entity recognition technology can also be used to match keywords in search statements with multiple preset character relationships to determine whether the keyword belongs to a semantic subject related to the character relationship.

[0352] When it is necessary to perform multiple processes including semantic subject filtering, spatiotemporal filtering, name filtering, and character relationship filtering on the candidate set recalled by the first branch, the execution order between these multiple processes can be set according to actual needs, and the embodiments of this application do not make specific limitations on this.

[0353] Exemplarily, as shown in FIG5 , visual media filtering includes:

[0354] 5101. Semantic subject filtering.

[0355] 5102. Space-time filtering.

[0356] 5103. Character relationship filtering.

[0357] 5104. Name filtering.

[0358] In the above 5101, generally, when searching, the user hopes to recall images that contain some of the visual content described in the search statement, such as: sky, puppy, child, etc. In other words, the semantic subject of the visual content in the search statement represents the visual focus of the user during the search.

[0359] However, when recalling photos, the first branch considers the degree of match between the visual content of visual media such as pictures and videos and the user's entire search statement, and does not fully consider the role of certain specific semantic subjects (that is, semantic subjects related to visual content) in the search statement in the mobile phone gallery search scenario, resulting in the recall of some inaccurate photos.

[0360] For example, when a user searches for "photos of the Great Wall taken during National Day," the semantic subject of the visual content includes "Great Wall." The first branch considers the degree of match between the entire search statement and the visual content of the image when performing the match. This can cause the first branch to recall images that contain the action of "taking photos" but not the "Great Wall." This is because the image contains the action of "taking photos" and the search statement contains the word "taking photos." In other words, there is a certain degree of similarity between the image's visual semantic vector and the search statement's sentence semantic vector, making the image likely to be recalled.

[0361] For example, when a user searches for "last year, my child celebrated his birthday holding a cake," the semantic entities of the visual content include "child," "birthday," and "cake." The first branch considers the degree of match between the entire search query and the image. This may cause the model to recall images of adults celebrating their birthdays holding cakes. This is because the visual semantic vector of the image of an adult celebrating his birthday holding a cake is very similar to the sentence semantic vector of "last year, my child celebrated his birthday holding a cake."

[0362] Therefore, in order to further improve the accuracy of searching for visual media on mobile phones, the "semantic subject" can be strengthened on the basis of the above-mentioned first branch recall of visual media, that is, the results of the above-mentioned first branch recall can be fine-tuned with the help of the degree of matching between the visual media and the "semantic subject".

[0363] Specifically, the "semantic subject filtering" in 5101 above may include the following steps:

[0364] 5101a. Determine the dimension to be matched based on the semantic subject related to the visual content in the search statement.

[0365] 5101b. Filter the M candidate visual media according to the matching degree of their visual contents in the to-be-matched dimension.

[0366] In one example, in 5101a above, the semantic subject related to the visual content is used as the dimension to be matched. If there are multiple semantic subjects related to the visual content in the search statement, each of these multiple semantic subjects is used as a different dimension to be matched, thereby obtaining multiple dimensions to be matched. The number of the multiple dimensions to be matched is the same as the number of the multiple semantic subjects related to the visual content.

[0367] For example, the search sentence is "Last year, my child celebrated his birthday with a cake", where the semantic subjects related to the visual content are "child", "birthday" and "cake". Then, the corresponding three dimensions to be matched are: "child", "birthday" and "cake".

[0368] In practical applications, in addition to using the semantic subject as a dimension to be matched, the search statement itself can also be used as a dimension to be matched. In this way, when performing semantic subject filtering, the matching situation between the visual content of the candidate visual media and the entire search statement can also be taken into account to improve the rationality of the filtering. Specifically, the semantic subject and the search statement related to the visual content can be used as different dimensions to be matched, respectively, to obtain multiple dimensions to be matched. When there are multiple semantic subjects related to the visual content in the search statement, these multiple semantic subjects and the search statement are used as different dimensions to be matched. The number of the multiple dimensions to be matched is one more than the number of the multiple semantic subjects related to the visual content.

[0369] For example, the search sentence is "Last year, my child celebrated his birthday holding a cake", where the semantic subjects of the visual content are "child", "birthday" and "cake". Then, the corresponding four dimensions to be matched are: "child", "birthday", "cake" and "Last year, my child celebrated his birthday holding a cake".

[0370] In the above 5101b, the degree of matching of the visual content of the candidate visual media on the dimension to be matched refers to the degree of matching of the visual content of the candidate visual media with the dimension to be matched. The vector similarity between the visual semantic vector of the candidate visual media and the semantic vector to be matched of the dimension to be matched can be determined; wherein, when the dimension to be matched is the semantic subject of the relevant visual content, the semantic vector to be matched of the dimension to be matched is the subject semantic vector of the semantic subject; when the dimension to be matched is a search sentence, the semantic vector to be matched of the dimension to be matched is the sentence semantic vector of the search sentence; based on the vector similarity, the degree of matching of the visual content of the candidate visual media on the dimension to be matched is determined. The matching degree is positively correlated with the vector similarity.

[0371] When the number of dimensions to be matched is one, candidate visual media with a matching degree less than or equal to a preset matching degree threshold may be filtered out based on the matching degrees of the visual contents of the plurality of candidate visual media on the dimensions to be matched.

[0372] For example, if the search statement is "photos of the sky taken during National Day", the only semantic subject related to the visual content is "sky". Then, the recalled pictures that contain the action of "taking" but not "sky" have a low match with "sky" and will be filtered out.

[0373] When there are multiple dimensions to be matched, for each candidate visual media, the comprehensive matching degree of the candidate visual media is determined based on the matching degree of the visual content of the candidate visual media on the multiple dimensions to be matched; and the multiple candidate visual media are filtered based on the comprehensive matching degree of each candidate visual media. Specifically, from the M candidate visual media, multiple first visual media with a comprehensive matching degree greater than or equal to a preset matching degree threshold (that is, meeting the preset requirements) are determined, which is equivalent to filtering out the candidate visual media with a comprehensive matching degree less than the preset matching degree threshold; or, the M candidate visual media are sorted from high to low according to the comprehensive matching degree, and the top Z′ (Z′≥1) candidate visual media (that is, meeting the preset requirements) are taken as multiple first visual media, which is equivalent to filtering out the (MZ′) candidate visual media with a lower ranking.

[0374] In an optional embodiment, any of the following three methods may be used to determine the comprehensive matching degree of the candidate visual media:

[0375] Method 1: summing up the matching degrees of the visual content of the candidate visual media in multiple dimensions to be matched to obtain the comprehensive matching degree of the candidate visual media.

[0376] Method 2: performing weighted summation on the matching degrees of the visual content of the candidate visual media in multiple dimensions to be matched, so as to obtain the comprehensive matching degree of the candidate visual media.

[0377] The weights of the multiple dimensions to be matched can be configured in advance by the user.

[0378] Method three: using a machine learning model to determine the comprehensive matching degree of the candidate visual media based on the matching degree of the visual content of the candidate visual media in multiple dimensions to be matched.

[0379] Among them, the machine learning model needs to be trained based on the data set, and the purpose of its training is essentially to learn the weights of each dimension to be matched.

[0380] In the above method 1, the contribution of the matching degrees in different dimensions to the comprehensive matching degree is not distinguished, which may result in the final calculated comprehensive matching degree values ​​of multiple candidate visual media being close or equal, and thus making it impossible to screen multiple candidate visual media.

[0381] For example, assume that the multiple candidate visual media include picture 1, picture 2, and picture 3, and the multiple dimensions to be matched include dimension A, dimension B, and dimension C. The matching degree of the visual content of each picture in each dimension is calculated respectively, and the results are shown in Table 1.

[0382] Table 1:

[0383] If calculated according to method 1, the comprehensive matching scores of pictures 1, 2, and 3 are all 1.2, which will make it impossible to filter the three pictures.

[0384] In the second method described above, when there are too many semantic subjects related to the visual content that need to be paid attention to (that is, the number of preset multiple tags is too large), it is difficult to accurately configure the weights of different dimensions.

[0385] In the third method above, the construction of the dataset is inseparable from user data. However, user data is private and users do not want their data to be reported to the cloud.

[0386] In an optional implementation, in order to solve the above problem, the following steps can be used to determine the comprehensive matching degree:

[0387] S51: Determine the weights of the N dimensions to be matched.

[0388] The weight of the jth dimension to be matched is positively correlated with the degree of variation of the matching degree of the visual contents of the M candidate visual media files on the jth dimension to be matched; j is an integer, and the value of j ranges from 1 to N.

[0389] S52 , performing weighted summation of the matching degrees of the visual content of the i th candidate visual media file in each of the N dimensions to be matched according to the weights of the N dimensions to be matched, to obtain a comprehensive matching degree of the i th candidate visual media file.

[0390] Wherein, i is an integer, and the value of i ranges from 1 to M.

[0391] For example, consider the search phrase "Last year, my child held a cake for his birthday." The semantic entities associated with the visual content include: child, birthday, and cake. If multiple recalled photos all contain cake, then the variation in the matching degree of the visual content of these recalled photos on the dimension of "cake" will be relatively small, and the corresponding weight for this dimension will be relatively small. However, if some recalled photos contain children while others do not, then the variation in the matching degree of the visual content of these recalled photos on the dimension of "child" will be relatively large, and the corresponding weight for this dimension will be relatively large.

[0392] In one example, for each dimension to be matched, the degree of variation in the matching degrees of the visual content of the candidate visual media along that dimension is determined based on the information entropy of the matching degrees of the visual content of the candidate visual media along that dimension; the degree of variation is inversely proportional to the information entropy. The calculation method of information entropy will be described in detail in the following embodiments.

[0393] In this embodiment, information entropy is used to measure the variation of the matching degree under each dimension to be matched, and based on this, the weight corresponding to the dimension to be matched is determined.

[0394] In order to ensure that the matching degree of the visual content of the candidate visual media on different dimensions to be matched has a uniform dimension, a normalization processing step can be performed. Specifically, for each candidate visual media, the initial matching degree of the candidate visual media on the dimension to be matched is determined based on the vector similarity between the visual semantic vector of the candidate visual media and the semantic vector to be matched of the dimension to be matched. Exemplarily, the vector similarity between the visual semantic vector of the candidate visual media and the semantic vector to be matched of the dimension to be matched can be used as the initial matching degree of the candidate visual media on the dimension to be matched. The initial matching degrees of the visual contents of multiple candidate visual media on the dimension to be matched are normalized to obtain the matching degrees of the visual contents of multiple candidate visual media on the dimension to be matched.

[0395] The following is a detailed introduction to the normalization process and the calculation process of weights and comprehensive matching degrees:

[0396] Assuming there are m candidate visual media and n dimensions to be matched, the initial matching degree of the visual content of the m candidate visual media in each of the n dimensions to be matched can be regarded as a data matrix:

[0397] X=(x ij ) m*n , i=1,2,…,m; j=1,2,…,n (5)

[0398] Among them, x ij is the initial matching degree of the visual content of the i-th candidate visual media on the j-th dimension to be matched.

[0399] Continuing with the above example, the multiple candidate visual media include Picture 1, Picture 2, and Picture 3, and the multiple dimensions to be matched include Dimension A, Dimension B, and Dimension C, that is, m is 3 and n is 3.

[0400] Step 1: Normalize the above data matrix.

[0401] The normalized matrix is:

[0402] R=(r ij ) m*n , i=1,2,…,m; j=1,2,…,n (6)

[0403] in,

[0404] Among them, max(x j ) refers to the maximum value of the initial matching degree of the visual content of multiple candidate visual media on the jth dimension to be matched; min(x j ) refers to the minimum value of the initial matching degree of the visual content of multiple candidate visual media on the j-th dimension to be matched.

[0405] In practical applications, other normalization methods may also be used, and the embodiments of the present application do not specifically limit this.

[0406] For example, the results obtained after normalizing the example data in Table 1 are shown in Table 2.

[0407] Table 2:

[0408] Step 2: Calculate the information entropy corresponding to each dimension to be matched.

[0409] Formula (8) and formula (9) can be used to calculate:

[0410] in,

[0411] Among them, e j Refers to the information entropy corresponding to the j-th dimension to be matched.

[0412] Among them, since the domain of ln(x) function is x>0. In actual calculation, in order to avoid ln(p ij ) ij When the value is 0, the ln(p ij) is replaced by ln(p ij +α), where α<<0.001.

[0413] The smaller the information entropy corresponding to the dimension to be matched, the greater the variation in the matching degree of the visual content of multiple candidate visual media on the dimension to be matched, and the greater the amount of information provided. It can be considered that the role of the dimension to be matched in the comprehensive evaluation is also greater.

[0414] For example, for the example data in Table 2 above, the information entropy calculated according to the above formula (8) and formula (9) is as shown in Table 3:

[0415] Table 3:

[0416] Step 3: Calculate the weight corresponding to each dimension to be matched.

[0417] Formula (10) can be used to calculate the weight corresponding to each dimension to be matched:

[0418] Among them, d j Refers to the weight corresponding to the j-th dimension to be matched.

[0419] It can be seen that the above formula (10) is a monotonically decreasing function of information entropy. In this embodiment, the smaller the information entropy, the greater the degree of variation; the greater the degree of variation, the greater the weight. In other words, the weight is negatively correlated with the information entropy.

[0420] In order to ensure that the sum of the weights corresponding to multiple dimensions to be matched is 1, the following calculation can be performed using formula (11) to obtain the final weight w j :

[0421] Among them, w j Refers to the final weight corresponding to the j-th dimension to be matched.

[0422] In an optional embodiment, other monotonically decreasing functions may be used to calculate the above weights, which is not specifically limited in the embodiment of the present application.

[0423] For example, for the example data in Table 3 above, the weights of different dimensions to be matched can be calculated according to the above formulas (10) and (11), as shown in Table 4:

[0424] Table 4:

[0425] It should be noted that, because rounding is introduced in the process of calculating information entropy, the sum of the three dimensions in Table 4 above is not 1.

[0426] Step 4: Weighted summation.

[0427] The normalized matching degree of any candidate visual media is weighted and summed to obtain the comprehensive matching degree of the candidate visual media. Specifically, it can be calculated using formula (12):

[0428] Among them, s i Refers to the comprehensive matching degree of the i-th visual media.

[0429] For example, for the example data in Table 2 and Table 4 above, the above formula (12) is used to calculate the comprehensive matching degree of different pictures, as shown in Table 5:

[0430] Table 5:

[0431] In 5102 above, the first branch only considers the visual semantics of the visual media when recalling them, without considering attributes such as time and location. Therefore, the visual media retrieved by the first branch need to be filtered by time and location. Specifically, when a user performs a semantic search and the search query includes time information, images that do not meet the time constraints need to be filtered out.

[0432] However, users' descriptions of time information are rich and varied, often fuzzy. For example, when a user searches for "sky photographed in the afternoon," the term "afternoon" in the search statement lacks a clear, standardized definition. When users use fuzzy time expressions in their search statements, the time window used for time filtering will affect the user experience.

[0433] For example: the user starts traveling on September 30, 2022, and arrives home on October 9, 2022; during this period, the user takes a lot of photos. One day in 2023, the user wants to view the beautiful scenery taken on this trip, so the user is likely to enter the search sentence "scenery taken during last year's National Day holiday". If the time window used for time filtering is set to "October 1, 2022 to October 7, 2022" directly based on "last year's National Day" in the search sentence, then the scenery photos taken by the user on September 30, 2022, October 8, 2022, and October 9, 2022 will be filtered out, which obviously does not meet the user's expectations. In order to improve the rationality of time filtering, an embodiment of the present application provides a new time filtering method. Specifically, a clustering algorithm is used to cluster multiple candidate visual media according to their respective acquisition times to obtain K (K≥1) cluster clusters.

[0434] In this way, images of the same series that were collected at a relatively close time can be classified into the same cluster.

[0435] Continuing with the above example, the natural scenery pictures taken by the user on September 30, 2022 and the natural scenery pictures taken by the user on October 1, 2022 were taken relatively close in time and are classified into the same cluster through the above clustering algorithm.

[0436] The above clustering algorithms may include but are not limited to: K-Means clustering algorithm, mean shift clustering algorithm and density-based clustering algorithm.

[0437] Taking density-based clustering as an example, in semantic search scenarios involving time information, it's unreasonable to set fixed parameters for the first parameter ε and the second parameter MinPts. For example, when a user searches for "photos of a trip in 2022," the time range is a full year; while when a user searches for "photos of a trip in the morning," the time range is a few hours. The first parameter ε and the second parameter MinPts should not be set identically in these two cases. To improve clustering rationality, the following steps can be used to determine the first parameter ε and the second parameter MinPts:

[0438] S1021. Determine a first parameter involved in a density-based clustering algorithm based on a time search range included in the search statement.

[0439] The first parameter is positively correlated with the duration corresponding to the time search range.

[0440] The time search range can be determined based on the time-related semantic subject in the search statement. Specifically, the time range corresponding to the time-related semantic subject (i.e., the time search range) can be returned through a mapping table. The mapping table can be constructed in advance as needed, and the specific form of the embodiment of the present application is not specifically limited to this.

[0441] For example, the semantic subject of time extracted from the search statement "photos taken in spring" is "spring", and the mapping table is queried to obtain the corresponding time range: February 1st-May 30th; the semantic subject of time extracted from the search statement "the sky taken in the morning" is "morning", and the mapping table is queried to obtain the corresponding event range: 7:00-12:00; the semantic subject of time extracted from the search statement "photos of traveling during National Day" is "National Day", and the mapping table is queried to obtain the corresponding time range: October 1st-October 7th; the subject of time extracted from the search statement "the sky taken in Beijing this year" is "this year", and the mapping table is queried to obtain the corresponding time range: "January 1, 2023 to December 31, 2023".

[0442] The first parameter can be determined according to the duration of the time search range. For example, the start time of the time search range (which can be understood as the start timestamp) is Tstart and the end time (which can be understood as the end timestamp) is T end , the duration of the time search range is: T end -T start , the first parameter can be calculated using the following formula:

[0443] ε=α*(T end -T start ) (13)

[0444] Here, α is a coefficient that can be adjusted manually, and its size can be set according to actual needs. This application does not make any specific restrictions on this.

[0445] S1022. Determine a second parameter involved in the density-based clustering algorithm based on a ratio of the number of multiple candidate visual media to the time search range.

[0446] The second parameter can be calculated using the following formula:

[0447] Among them, N is the total number of visual media in the candidate set whose collection time is within the time search range, N≥1; among them, n is the number of times the time search range is repeated from T1 to T2. T1 is the collection time of the earliest collected visual media among the multiple visual media stored by the mobile phone; T2 is the collection time of the latest collected visual media among the multiple visual media stored by the mobile phone. Exemplarily, the search statement is "the sky photographed on National Day", and its time search range is: October 1 to October 7, which is 7 days long; the collection time of the earliest collected visual media stored in the mobile phone is August 1, 2020; the collection time of the latest collected visual media stored in the mobile phone is October 20, 2023; then, from August 1, 2020 to October 20, 2023, the time search range from October 1 to October 7 is repeated 4 times (that is, once a year).

[0448] Among them, (T end -T start )*n can be understood as the total duration corresponding to the time search range.

[0449] Wherein, β is a coefficient that can be adjusted manually, and its size can be set according to actual needs, and this application does not make any specific limitation on this.

[0450] In this embodiment, the first parameter ε and the second parameter MinPts involved in the density-based clustering algorithm are dynamically adjusted according to the duration limited by the time search range contained in the search statement, which can ensure the rationality of clustering and thus improve the accuracy of the final search results.

[0451] When the search statement contains a time search range, determine the collection time range corresponding to each of K (K ≥ 1) clusters; filter out clusters whose collection time range does not overlap with the time search range, and retain clusters whose collection time range does overlap with the time search range.

[0452] The overlap between the acquisition time range and the time search range may be partial overlap or full overlap. Regardless of partial overlap or full overlap, both are considered to have overlapping parts.

[0453] Thus, G (G ≥ 1) clusters are selected from the K clusters. The earliest visual media collected in these G clusters is collected before the start time of the time search range, and / or the latest visual media collected in these G clusters is collected after the end time of the time search range. Note: There is no intersection between the G clusters, and the collection time ranges of the G clusters do not overlap.

[0454] The collection time range corresponding to the cluster is the collection time T of the earliest collected visual media in the cluster. min T1 to the acquisition time T of the latest visual media in the cluster max , that is: [T min , T max ].

[0455] Specifically, the time search range is [T start , T end ], the collection time range corresponding to the cluster is [T min , T max ], when [T min , T max ] and [T start , T end ]When there is an overlap, the cluster is retained; when [T min , T max ] and [T start , T end If there is no overlap, the cluster is filtered. For example, if the time search range is from October 1, 2022 to October 7, 2022, and the collection time range of the cluster is from September 30, 2022 to October 1, 2022, and there is an overlap between the two ranges (that is, October 1, 2022), the cluster is retained.

[0456] Continuing with the above example, a natural scenery image taken by a user on September 30, 2022, and another natural scenery image taken by the user on October 1, 2022, were taken close together and are therefore grouped into the same cluster using the above clustering algorithm. Assuming this cluster contains both images, the collection time range for this cluster is September 30, 2022, to October 1, 2022. This overlaps with the time search range of October 1, 2022, to October 7, 2022. Therefore, this cluster is retained, meaning that the natural scenery image taken by the user on September 30, 2022, will not be filtered out.

[0457] Similarly, a natural scenery image taken by a user on October 8, 2022, and another natural scenery image taken by a user on October 7, 2022, were taken close together and are therefore grouped together in the same cluster using the above clustering algorithm. Assuming this cluster contains both images, the collection time range for this cluster is from October 7, 2022, to October 8, 2022. This overlaps with the time search range of October 1, 2022, to October 7, 2022, so this cluster is retained. In other words, the natural scenery image taken by the user on October 8, 2022, will not be filtered out.

[0458] In this way, for the search sentence "scenery taken during last year's National Day holiday", the earliest collected visual media in the multiple clusters screened out was collected on September 30, 2022, and the latest collected visual media was collected on October 8, 2022.

[0459] It can be seen that the time filtering method provided in the embodiment of the present application can ensure that photos of the same series with relatively close collection time are displayed to the user, ensure the consistency of search results, and improve the user's search experience.

[0460] In addition, when the search statement also contains semantic subjects related to locations, the G clusters selected can continue to be filtered by location. Specifically, visual media in the cluster whose collection locations do not match the semantic subjects related to locations in the search statement can be filtered out. For example, the collection location "Prince Gong's Mansion, Xicheng District, Beijing" and the semantic subject "Sujiatuo Town, Haidian District, Beijing" can be considered a match (belonging to the city level match); the collection location "Tsinghua University, Haidian District, Beijing" and the semantic subject "Sujiatuo Town, Haidian District, Beijing" can be considered a match (belonging to the district level match); the collection location "Shanghai" and the semantic subject "Beijing" can be considered a mismatch.

[0461] In the above embodiments, clustering is performed first, followed by time filtering, and finally location filtering. Of course, in actual applications, location filtering can also be performed first, followed by clustering, and finally time filtering. The specific execution order of these three steps can be set according to actual needs, and the embodiments of this application do not make specific limitations in this regard.

[0462] 5103. Filtering of relationship between people.

[0463] When the search statement contains a semantic subject related to the relationship between people, visual media in each clustering cluster whose relationship attribute between people does not match this semantic subject is filtered out. Exemplarily, if the search statement contains "good friend" and the name attribute of Picture B is "colleague", and these two do not match, then Picture B is filtered out.

[0464] 5104. Filtering of names.

[0465] When the search statement includes a semantic subject related to a name, visual media in each clustering cluster whose name attribute does not match this semantic subject is filtered out. Exemplarily, if the search statement contains "Zhang San" and the name attribute of Picture A is "Li Si", and these two do not match, then Picture A is filtered out.

[0466] It should be supplemented that the semantic subject related to a name and the semantic subject related to the relationship between people in the search statement can also be identified through named entity recognition technology. The name attribute and the relationship attribute between people of visual media are manually added by the user for visual media in advance.

[0467] 511. Sorting.

[0468] For the candidate set obtained in step 508 above, the visual media in the candidate set can be sorted according to the acquisition time of the visual media in the candidate set, and the display order of the visual media is obtained. The acquisition time of the visual media with a higher display order is earlier than the acquisition time of the visual media with a lower display order. Subsequently, the mobile phone can display the candidate set according to the display order of the visual media in the candidate set.

[0469] For the filtered candidate set (including the above G clustering clusters) obtained in step 510 above, one of the following methods can be used for sorting:

[0470] Method 1: Sort the visual media in the G clustering clusters according to the target matching degree of the visual media in the G clustering clusters from high to low, and obtain the display order of the visual media in the G clustering clusters. This target matching degree can be the matching degree of the visual content of the visual media and the search statement in the above text or the comprehensive matching degree in the above text (the specific calculation method can refer to the corresponding content in the above embodiments). Subsequently, the mobile phone can display the visual media in the G clustering clusters according to the display order of the visual media in the G clustering clusters.

[0471] Method 2: Sort the G clusters according to the start time of their collection time ranges to obtain the display order of the G clusters (i.e., inter-cluster sorting). For example, as shown in interface 601 of FIG6 , the collection time range corresponding to cluster A is from September 30, 2022 to October 1, 2022; the collection time range corresponding to cluster B is from October 3, 2022 to October 5, 2022; then, the display order of cluster A is prior to the display order of cluster B. For each cluster, sort the visual media within the cluster from high to low according to the target matching degree of the visual media within the cluster to obtain the display order between the visual media within the cluster (intra-cluster sorting); or, for each cluster, sort the visual media within the cluster according to the collection time of the visual media within the cluster to obtain the display order of the visual media within the cluster. Subsequently, the mobile phone displays the visual media in the G clusters according to the inter-cluster sorting and intra-cluster sorting.

[0472] It should be noted that in order to better protect user privacy and security and meet the principle of minimizing user data, user data should be avoided from being reported to the cloud side as much as possible, and the entire search process mentioned above should be completed on the terminal side.

[0473] In addition, the present application provides an electronic device comprising: a memory, a processor and a display, wherein the memory is used to store programs; the processor is coupled to the memory and the display, and is used to execute the programs stored in the memory to implement the above-mentioned visual media search method.

[0474] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer, can implement one or more steps in any of the above-mentioned visual media search methods.

[0475] The computer readable storage medium may be a non-transitory computer readable storage medium, for example, a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, and the like.

[0476] Another embodiment of the present application further provides a computer program product comprising instructions, which, when executed by a computer, can implement one or more steps in any of the above methods.

[0477] Among them, the electronic device, computer-readable storage medium, and computer program product provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.

[0478] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0479] Units described as separate components may or may not be physically separate, and components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0480] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0481] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0482] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A visual media search method, applicable to electronic devices, characterized in that: include: Display the first interface; The first interface includes a search box; receiving a search operation on a search statement input into the search box, wherein the search statement includes a time search range; Using a clustering algorithm, clustering the plurality of first visual media files according to their acquisition time to obtain K clusters, wherein the acquisition time ranges of the K clusters are different; wherein K ≥ 1 and is an integer; Determine a target cluster according to whether the acquisition time range partially overlaps with the time search range; Determining search results of the search statement according to the target cluster; The search results are displayed.

2. The method according to claim 1, characterized in that: The start time of the collection time range of each cluster is the collection time of the earliest collected visual media file in the cluster, and the end time is the collection time of the latest collected visual media file in the cluster; There is no intersection between the collection time ranges of the K clusters.

3. The method according to claim 2, characterized in that The start time of the time search range is the first time point, and the end time is the second time point; Determining a target cluster according to whether the acquisition time range partially overlaps with the time search range includes: When the start time of the acquisition time range of the kth cluster among the K clusters is between the first time point and the second time point and the end time is not between the first time point and the second time point, determining that the acquisition time range of the kth cluster partially overlaps with the time search range; When the end time of the acquisition time range of the kth cluster is between the first time point and the second time point and the start time is not between the first time point and the second time point, determining that the acquisition time range of the kth cluster partially overlaps with the time search range; When the start time of the acquisition time range of the kth cluster is less than the first time point and the end time is greater than the second time point, it is determined that the acquisition time range of the kth cluster partially overlaps with the time search range; k is an integer, 1≤k≤K; A cluster whose collection time range partially overlaps with the time search range among the K clusters is determined as a target cluster.

4. The method according to claim 2, characterized in that: The search results include a plurality of the target clusters; The method further comprises: Determining a display order of the plurality of target clusters according to a chronological order of start times of the acquisition time ranges of the plurality of target clusters; and displaying the search results, including: The plurality of target clusters are displayed in a display order of the plurality of target clusters.

5. The method according to any one of claims 1 to 4, characterized in that Also includes: Performing semantic understanding on the search sentence to obtain a semantic vector of the first sentence; Obtaining visual semantic vectors of multiple visual media files; The visual semantic vector of each visual media file is obtained by using a natural picture understanding model to perform semantic understanding on the image or image frame of the visual media file; Determine, from the plurality of visual media files, M candidate visual media files whose visual semantic vectors match the first sentence semantic vector; wherein M is an integer greater than 1; The plurality of first visual media files are determined according to the M candidate visual media files.

6. The method according to claim 5, characterized in that Using a clustering algorithm, clustering the plurality of first visual media files according to their acquisition time to obtain K clusters, including: If the proportion of semantics related to the visual content in the search statement is greater than or equal to a preset proportion threshold, clustering the multiple first visual media files according to the acquisition time of the multiple first visual media files using a clustering algorithm to obtain K clusters; the visual content is data that needs to be acquired by a natural image understanding model; The method further comprises: obtaining properties of the plurality of visual media files; From the plurality of visual media files, determining a plurality of second visual media files having attributes matching the semantic subject in the search statement; If the proportion of semantics related to visual content in the search statement is less than the preset proportion threshold, the search results of the search statement are determined according to the multiple second visual media.

7. The method according to claim 6, characterized in that Also includes: Removing semantic entities irrelevant to the visual content from the search statement to obtain a rewritten search statement; Performing semantic understanding on the rewritten search sentence to obtain a second sentence semantic vector; Determining, from the plurality of visual media files, a plurality of reference visual media whose visual semantic vectors match the second sentence semantic vector; Determining first representative visual semantic vectors corresponding to the plurality of first visual media files and second representative visual semantic vectors corresponding to the plurality of reference visual media; Determining a semantic proportion related to the visual content in the search statement according to a vector similarity between the first difference vector and the second difference vector; The first difference vector represents the difference between the first sentence semantic vector and the first representative visual semantic vector; The second difference vector represents a difference between the second sentence semantic vector and the second representative visual semantic vector.

8. The method according to claim 7, characterized in that The semantic subject irrelevant to the visual content in the search statement is removed to obtain a rewritten search statement, including: The time-related semantic subject and / or the location-related semantic subject in the search statement are removed to obtain a rewritten search statement.

9. The method according to claim 5, characterized in that Determining the plurality of first visual media files according to the M candidate visual media files comprises: Determine, according to the semantic subject related to the visual content in the search statement, N dimensions to be matched corresponding to the search statement, where N is greater than 1 and is an integer; Obtaining weights of the N dimensions to be matched; wherein the weight of the jth dimension to be matched is positively correlated with the variation degree of matching of the visual contents of the M candidate visual media files on the jth dimension to be matched; j is an integer, and the values ​​of j range from 1 to N in sequence; According to the weights of the N dimensions to be matched, weighted summation is performed on the matching degrees of the visual content of the i-th candidate visual media file in each of the N dimensions to be matched, so as to obtain a comprehensive matching degree of the i-th candidate visual media file; i is an integer, and the values ​​of i range from 1 to M in sequence; The plurality of first visual media files whose comprehensive matching degrees meet preset requirements are determined from the M candidate visual media files.

10. The method according to claim 9, characterized in that Also includes: Matching the search term in the search statement with a plurality of preset tags to determine whether the search term belongs to a semantic subject related to the tag; The tags are used to describe the visual content; The semantic subject related to the tag in the search sentence is determined as the semantic subject related to the visual content in the search sentence.

11. The method according to any one of claims 1 to 4, characterized in that The clustering algorithm includes: a density-based clustering algorithm; The method further comprises: According to the time search range, a first parameter involved in the density-based clustering algorithm is determined; the first parameter is used to describe the neighborhood radius of the data point; the first parameter is positively correlated with the duration corresponding to the time search range.

12. The method according to claim 11, characterized in that Also includes: determining the number of visual media files among the plurality of first visual media files whose acquisition time is within the time search range; A second parameter involved in the density-based clustering algorithm is determined based on the ratio of the number of visual media files to the total duration corresponding to the time search range; the second parameter is used to describe the minimum number of data points in the neighborhood of a data point.

13. The method according to any one of claims 1 to 4, characterized in that Determining search results of the search statement according to the target cluster includes: If the search statement includes a semantic subject related to a location, filtering the visual media files in the target cluster according to the semantic subject related to the location and the collection location attributes of the visual media files in the target cluster; If the search statement includes a semantic subject related to character relationships, filtering the visual media files in the target cluster according to the semantic subject related to character relationships and character relationship attributes of the visual media files in the target cluster; and / or If the search statement includes a semantic subject related to a person's name, the visual media files in the target cluster are filtered according to the semantic subject related to the person's name and the person's name attribute of the visual media files in the target cluster.

14. An electronic device, characterized in that: include: A memory, a processor, and a display, wherein: The memory is used to store programs; The display is used to display the search page; The processor is coupled to the memory and the display, and is configured to execute the program stored in the memory to implement the method according to any one of claims 1 to 13.

15. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a computer, the method according to any one of claims 1 to 13 can be implemented.

16. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Intelligent automated assistant for delivering content from user experiences

    CN110457000A

  • Visual media personalized search method and device

    CN113641857A

  • Interactive Photo Annotation Based on Face Clustering

    US20080298766A1

  • Enhanced image-search using contextual tags

    US20210149947A1

  • Search method and electronic device

    WO2023029993A1