Visual media search method and electronic device
By combining the base model and the fine-tuning model, candidate visual media are screened and confirmed, solving the problem of inaccurate visual media retrieval on mobile phones under complex search statements, and achieving highly accurate visual media search.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2023-11-23
- Publication Date
- 2026-05-26
AI Technical Summary
Existing mobile phones struggle to accurately retrieve visual media when processing complex search queries, resulting in inaccurate search results.
By combining the base model and the fine-tuning model, candidate visual media with similarity greater than a threshold are initially screened, and the target visual media are determined by secondary confirmation based on image feature vectors, text feature vectors and label information.
It improves the accuracy of visual media search, ensuring that target visual media can be accurately retrieved regardless of the complexity of the search query, thus meeting user needs.
Smart Images

Figure CN120067394B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a visual media search method and electronic device. Background Technology
[0002] With the development of electronic devices (such as mobile phones), the camera function of mobile phones has also developed rapidly. More and more users are using their mobile phones to take photos and videos and store them in their phone's gallery. In addition, users can also store downloaded pictures and screenshots in their phone's gallery.
[0003] When users want to search for visual media (such as photos and videos), they can enter a search query on their mobile phone, such as "photos taken on September 1st." The phone will respond to the user's search query, retrieve photos taken on September 1st, and obtain the corresponding search results. However, mobile phones have limited ability to understand search queries, and when the search query is complex, the phone may not be able to accurately obtain the corresponding search results. Summary of the Invention
[0004] In view of this, this application provides a visual media search method and electronic device to improve the accuracy of search results.
[0005] Firstly, this application provides a visual media search method applied to an electronic device. The electronic device can display a first interface, which includes a search box. The electronic device can receive a search query entered by a user in the search box. Then, the electronic device can determine candidate visual media, which represent visual media on the electronic device whose similarity to the search query is greater than a first threshold. Subsequently, for each candidate visual media, the electronic device can determine a matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the search query.
[0006] Then, the electronic device can determine the target visual media from the candidate visual media based on the matching values corresponding to each candidate visual media. The electronic device then displays the search results, which correspond to the target visual media.
[0007] In this application, after receiving a user's search query about visual media, the electronic device indicates that it needs to search for visual media that match the search query. The electronic device can select visual media with a similarity greater than a first threshold as candidate visual media, thus achieving an initial retrieval of visual media. Subsequently, the electronic device can perform a secondary verification of the candidate visual media. That is, based on the image information of the candidate visual media and the information of the search query, the electronic device can determine the matching value corresponding to the candidate visual media. This allows the electronic device to further determine whether the candidate visual media matches the search query based on the matching value, thereby determining whether the candidate visual media can be used as the target visual media in the search results, improving the accuracy of target visual media determination, and thus improving the accuracy of search result determination. Furthermore, by determining whether a visual media is the target visual media through the similarity between the visual media and the search query, as well as based on the image information of the visual media and the information of the search query, accurate searching of target visual media can be achieved regardless of the complexity of the search query, satisfying the user's search needs.
[0008] In one possible design approach, the image information of the candidate visual media includes at least one of the following: the similarity between the candidate visual media and the search statement, the image feature vector of the candidate visual media, and the first label of the candidate visual media.
[0009] The information of the search statement includes at least one of the following: the search statement, the text feature vector of the search statement, the second tags included in the search statement, and the word segmentation of the search statement.
[0010] The first label of the candidate visual media indicates the category to which the visual content of the candidate visual media belongs.
[0011] The second tag for the search query is obtained by mapping the word segmentation of the search query. Specifically, for each word segment of visual media, the electronic device can determine whether the word exists in the preset tag table. If the word exists in the preset tag table, the tag corresponding to the word in the preset tag table is used as the second tag.
[0012] Optionally, the similarity between the candidate visual media and the search statement may include the similarity between the candidate visual media and the search statement corresponding to the base model (that is, the similarity between the candidate visual media and the search statement determined by the base model) and / or the similarity between the candidate visual media and the search statement corresponding to the fine-tuning model (that is, the similarity between the candidate visual media and the search statement determined by the fine-tuning model).
[0013] The image feature vectors of the aforementioned candidate visual media may include the image feature vectors of the candidate visual media corresponding to the base model and / or the image feature vectors of the candidate visual media corresponding to the fine-tuning model.
[0014] The text feature vector of the search statement mentioned above may include the text feature vector of the search statement corresponding to the base model and / or the text feature vector of the search statement corresponding to the fine-tuning model.
[0015] The fine-tuned model described above is trained on the base model using a training set corresponding to a preset search scenario, and the fine-tuned model has a corresponding support range. When the search query falls within this support range, the fine-tuned model can accurately retrieve visual media that matches the search query.
[0016] In one possible design approach, the process of identifying candidate visual media described above may include:
[0017] For each visual medium on an electronic device, the electronic device can determine the similarity between the visual medium and the search query based on the image feature vector of the visual medium and the text feature vector of the search query.
[0018] If the similarity between the visual medium and the search query is greater than or equal to a first threshold, the electronic device determines that the visual medium is a candidate visual medium.
[0019] If the similarity between the visual medium and the search query is less than a first threshold, the electronic device determines that the visual medium is not a candidate visual medium.
[0020] In this application, the electronic device can calculate the vector similarity between the image feature vector of the visual media and the text feature vector of the search query. That is, based on the image feature vector of the visual media and the text feature vector of the search query, the similarity between the visual media and the search query is determined. Subsequently, the electronic device can use the similarity to determine whether the visual media is highly relevant to the search query, thereby identifying candidate visual media and ensuring the accuracy of the initial retrieval.
[0021] Optionally, if the similarity between the visual media corresponding to the base model and the search statement is greater than or equal to a first threshold, the electronic device determines that the visual media is a candidate visual media.
[0022] And / or,
[0023] If the similarity between the visual media corresponding to the fine-tuning model and the search statement is greater than or equal to a first threshold, the electronic device determines that the visual media is a candidate visual media.
[0024] In this application, electronic devices can determine the similarity between visual media and search queries using a base model and / or a fine-tuned model, thereby identifying candidate visual media based on this similarity. Because the base model has strong generalization ability, it can accurately determine the similarity between visual content in visual media and search queries in most search scenarios. The fine-tuned model can accurately identify the similarity between visual content in visual media and search queries in specific search scenarios.
[0025] In one possible design approach, the aforementioned candidate visual media includes visual media corresponding to the base model whose similarity to the search query is greater than or equal to a first threshold, and visual media corresponding to the fine-tuning model whose similarity to the search query is greater than or equal to the first threshold, thereby ensuring the comprehensiveness of the search results.
[0026] Optionally, the threshold for determining candidate visual media by comparing the similarity between the visual media corresponding to the fine-tuning model and the search statement may not be the first threshold, but the second threshold. That is, if the similarity between the visual media corresponding to the fine-tuning model and the search statement is greater than or equal to the second threshold, the electronic device can determine that the visual media is a candidate visual media.
[0027] In one possible design approach, the similarity between the visual media corresponding to the base model and the search query is determined based on the image feature vector and the text feature vector of the visual media corresponding to the base model.
[0028] The similarity between the visual media corresponding to the above fine-tuning model and the search statement is determined based on the image feature vector and the text feature vector of the visual media corresponding to the fine-tuning model.
[0029] In one possible design approach, the process of determining the matching value corresponding to the aforementioned candidate visual media may include:
[0030] If the first preset vocabulary does not include the search query, it indicates that the search query did not match the whitelist corresponding to the fine-tuning model. The electronic device can determine the matching value of the candidate visual media based on the first target matching value of the candidate visual media in the branch corresponding to the base model of the search query. The first target matching value includes at least one of a first matching value, a second matching value, a third matching value, and a fourth matching value; the first matching value is determined based on the matching degree between the candidate visual media and the search query, that is, the first matching value is determined by the binary classification module.
[0031] The second matching value is determined by comparing the similarity between the candidate visual media and the search statement with the first similarity threshold corresponding to the search statement; that is, the second matching value is determined by the dynamic threshold module.
[0032] The third matching value is determined based on whether the visual content in the candidate visual media includes the visual content corresponding to the search query. In other words, the third matching value is determined by the tag confirmation module.
[0033] The fourth matching value is determined by comparing the similarity between the candidate visual media and the search statement with the second similarity threshold corresponding to the search statement. In other words, the fourth matching value is determined by the whitelist threshold module.
[0034] If the first preset vocabulary includes the search query, it indicates that the search query hits the whitelist corresponding to the fine-tuning model. The electronic device can determine the matching value corresponding to the candidate visual media based on the second target matching value. The second target matching value includes the fifth matching value and / or the sixth matching value.
[0035] The fifth matching value is determined based on whether the visual content in the candidate visual media includes the visual content corresponding to the search query; that is, the fifth matching value is determined by the tag confirmation module.
[0036] The sixth matching value is determined by comparing the similarity between the candidate visual media and the search statement with the third similarity threshold corresponding to the search statement. In other words, the sixth matching value is determined by the whitelist threshold module.
[0037] In this application, the module used for secondary confirmation is determined by judging whether the search statement hits the first preset word list, that is, whether it hits the whitelist corresponding to the fine-tuning model, so as to ensure the accuracy of secondary confirmation.
[0038] In one possible design approach, the matching value mentioned above is a Boolean value. The process of determining the target visual media based on the matching value corresponding to the candidate visual media may include:
[0039] If the matching value corresponding to the above candidate visual media is true, the electronic device will use the candidate visual media as the target visual media.
[0040] If the matching value corresponding to the above candidate visual media is false, the electronic device determines that the candidate visual media is not the target visual media.
[0041] In one possible design approach, when the first preset thesaurus does not include the search query, the matching value corresponding to the aforementioned candidate visual media is the union of the matching values included in the first target matching value corresponding to the candidate visual media. Specifically, if the first target matching value corresponding to the candidate visual media includes true, the electronic device can determine that the matching value corresponding to the candidate visual media is true. If the first target matching value corresponding to the candidate visual media does not include true, that is, if all the matching values in the first target matching value are false, the electronic device can determine that the matching value corresponding to the candidate visual media is false.
[0042] Optionally, the aforementioned first target matching value may include a first matching value, a second matching value, a third matching value, and a fourth matching value. Accordingly, if the first matching value, the second matching value, the third matching value, and the fourth matching value are all false, the electronic device can determine that the matching value corresponding to the candidate visual media is false.
[0043] If the first matching value, the second matching value, the third matching value, or the fourth matching value is true, the electronic device can determine that the matching value corresponding to the candidate visual media is true.
[0044] In one possible design approach, when the first preset thesaurus includes the search query, the matching value corresponding to the candidate visual medium is the union of the matching values included in the second target matching value corresponding to the candidate visual medium. Specifically, if the second target matching value corresponding to the candidate visual medium includes "true", the electronic device can determine that the matching value corresponding to the candidate visual medium is "true".
[0045] If the second target matching value corresponding to the candidate visual medium does not include true, that is, if all the matching values in the second target matching value are false, the electronic device can determine that the matching value corresponding to the candidate visual medium is false.
[0046] Optionally, the aforementioned first target matching value may include a fifth matching value and a sixth matching value. Accordingly, if both the fifth and sixth matching values corresponding to the candidate visual medium are false, the electronic device can determine that the matching value corresponding to the candidate visual medium is false.
[0047] If the fifth and sixth matching values corresponding to the candidate visual media are true, the electronic device can determine that the matching value corresponding to the candidate visual media is true.
[0048] In one possible design approach, the aforementioned first target matching value may include a first matching value. The aforementioned image information of the candidate visual media includes the image feature vector of the candidate visual media, and the aforementioned search statement information includes the text feature vector of the search statement.
[0049] Accordingly, the process of determining the first matching value mentioned above may include:
[0050] Electronic devices can input the image feature vectors of each candidate visual medium and the text feature vectors of the search query into a binary classification model to obtain the first matching value corresponding to each candidate visual medium.
[0051] The binary classification model is used to determine the matching degree between the candidate visual media and the search statement based on the image feature vector of the candidate visual media and the text feature vector of the search statement. It also determines the first matching value corresponding to the candidate visual media based on the matching degree and the preset classification threshold. That is, if the matching degree between the candidate visual media and the search statement is greater than the preset classification threshold, the first matching value corresponding to the candidate visual media is determined to be true.
[0052] If the matching value between the candidate visual media and the search statement is less than or equal to a preset classification threshold, the first matching value corresponding to the candidate visual media is determined to be false.
[0053] In this application, electronic devices can use a binary classification model to determine whether the matching degree between candidate visual media and search terms is high, thereby determining whether candidate visual media matches search terms, realizing secondary determination of candidate visual media, and ensuring the accuracy of search results.
[0054] Optionally, in the case of the branch corresponding to the base model of the search statement, the image information of the candidate visual media can include the image information of the candidate visual media determined by the base model, and the information of the search statement can include the information of the search statement determined by the base model. Therefore, the image feature vector of the candidate visual media used by the binary classification model includes the image feature vector of the candidate visual media corresponding to the base model, and the text feature vector of the search statement includes the text feature vector of the search statement corresponding to the base model, ensuring the accuracy of determining the first matching value corresponding to the candidate visual media.
[0055] In one possible design, the first target matching value may include a second matching value. The image information of the candidate visual media includes the similarity between the candidate visual media and the search query. The information of the search query may include the search query itself or the word segmentation results of the search query, which include at least one word.
[0056] Accordingly, the process of determining the second matching value mentioned above may include:
[0057] Electronic devices can determine a first similarity threshold for a search query based on its length. The length of the search query can be determined by statistically analyzing the length of the search query itself or by analyzing the length of its word segments.
[0058] Then, for each candidate visual medium, the electronic device can determine whether the similarity between the candidate visual medium and the search statement is greater than or equal to a first similarity threshold.
[0059] If the similarity between the candidate visual medium and the search statement is greater than or equal to the first similarity threshold, the second matching value corresponding to the candidate visual medium is determined to be true;
[0060] If the similarity between the candidate visual medium and the search statement is less than the first similarity threshold, the second matching value corresponding to the candidate visual medium is determined to be false.
[0061] Among them, the first similarity threshold corresponding to the above search statement is positively correlated with the length of the search statement.
[0062] In this application, considering that the longer the search statement, the more visual content in the visual media needs to match the search statement, for example, if the search statement includes three semantic subjects, the more visual content in the visual media matches these three semantic subjects, the better the visual media matches the search statement. Therefore, the electronic device can use the length of the search statement to determine the first similarity threshold corresponding to the search statement. Then, the electronic device can compare the similarity between the candidate visual media and the search statement with the first similarity threshold to determine that the visual content in the candidate visual media has a high degree of matching with the search statement, thereby obtaining the second matching value corresponding to the candidate visual media. This ensures the accuracy of the second matching value determination, thus enabling accurate secondary confirmation of the candidate visual media.
[0063] Optionally, the similarity between the candidate visual media and the search statement used to determine the second matching value includes the similarity between the candidate visual media and the search statement corresponding to the base model.
[0064] In one possible design approach, the aforementioned first target matching value may include a third matching value. The image information of the candidate visual media includes the first tag of the candidate visual media; the information of the aforementioned search statement includes the word segmentation and second tag of the search statement;
[0065] Accordingly, the process for determining the third matching value mentioned above may include:
[0066] If the first tag of the candidate visual media includes a word segment of the search query, or if the second tag of the search query is included, then the third matching value corresponding to the candidate visual media is determined to be true.
[0067] If the first tag of the candidate visual media does not include the word segment of the search query, and the second tag of the search query is not included, then the third matching value corresponding to the candidate visual media is determined to be false.
[0068] Specifically, if the first tag of the candidate visual media includes at least one word of the search query, the electronic device can determine that the first tag of the candidate visual media includes the word of the search query. If the first tag of the candidate visual media does not include any of the words of the search query, the electronic device can determine that the first tag of the candidate visual media does not include any of the words of the search query. Similarly, if the first tag of the candidate visual media includes at least one second tag included in the search query, the electronic device can determine that the first tag of the candidate visual media includes the second tag of the search query. If the first tag of the candidate visual media does not include any of the second tags of the search query, the electronic device can determine that the first tag of the candidate visual media does not include any of the second tags of the search query.
[0069] Alternatively, to improve search accuracy, if the first tag of the candidate visual media includes each word of the search query, the electronic device can determine that the first tag of the candidate visual media includes the word of the search query. If the first tag of the candidate visual media does not include at least one word of the search query, the electronic device can determine that the first tag of the candidate visual media does not include the word of the search query. Similarly, if the first tag of the candidate visual media includes each of the second tags included in the search query, the electronic device can determine that the first tag of the candidate visual media includes the second tags of the search query. If the first tag of the candidate visual media does not include at least one second tag of the search query, the electronic device can determine that the first tag of the candidate visual media does not include the second tags of the search query.
[0070] In this application, the electronic device can determine whether the search statement matches the visual content of the candidate visual media by judging whether the search statement hits the first tag of the candidate visual media, thereby achieving secondary determination of the candidate visual media.
[0071] In one possible design approach, the aforementioned first target matching value may include a third matching value. The image information of the candidate visual media includes the first label of the candidate visual media; the information of the aforementioned search statement includes the word segmentation of the search statement.
[0072] Accordingly, the process for determining the third matching value mentioned above may include:
[0073] If the first tag of the candidate visual media includes the word segment of the search query, the third matching value corresponding to the candidate visual media is determined to be true; if the first tag of the candidate visual media does not include the word segment of the search query, the third matching value corresponding to the candidate visual media is determined to be false.
[0074] In one possible design approach, the aforementioned first target matching value may include a third matching value. The image information of the candidate visual media includes a first label for the candidate visual media; the information of the aforementioned search statement includes a second label for the search statement.
[0075] Accordingly, the process for determining the third matching value mentioned above may include:
[0076] If the first tag of the candidate visual media includes the second tag of the search statement, determine that the third matching value corresponding to the candidate visual media is true.
[0077] If the first tag of the candidate visual media does not include the second tag of the search query, the third matching value corresponding to the candidate visual media is determined to be false.
[0078] In one possible design approach, the aforementioned first target matching value may include a fourth matching value. The image information of the candidate visual media includes the similarity between the candidate visual media and the search statement, and the information of the search statement includes the search statement itself.
[0079] The process of determining the fourth matching value mentioned above may include:
[0080] When the second preset vocabulary includes a search statement, for each candidate visual medium, if the similarity between the candidate visual medium and the search statement is greater than or equal to the second similarity threshold, the electronic device can determine that the fourth matching value corresponding to the candidate visual medium is true; the second similarity threshold refers to the similarity threshold corresponding to the search statement in the second preset vocabulary.
[0081] If the similarity between the candidate visual medium and the search statement is less than the second similarity threshold, the electronic device can determine that the fourth matching value corresponding to the candidate visual medium is false.
[0082] In this application, the electronic device can determine whether a search statement matches a second preset vocabulary. When the search statement matches the second preset vocabulary, the electronic device can determine a second similarity threshold corresponding to the search statement from the second preset vocabulary, thus accurately determining the similarity threshold. Subsequently, the electronic device can compare the similarity between the candidate visual media and the search statement with the relationship between the second similarity threshold to determine whether the similarity between the candidate visual media and the search statement is sufficiently high, thus achieving accurate secondary confirmation of the candidate visual media.
[0083] Optionally, the similarity between the candidate visual media and the search statement used by the electronic device to determine the fourth matching value can be the similarity between the candidate visual media and the search statement corresponding to the base model.
[0084] In one possible design approach, the second target matching value may include a fifth matching value. The image information of the candidate visual media includes the first tag of the candidate visual media; the information of the search statement includes the word segmentation and second tag of the search statement.
[0085] Accordingly, the process of determining the fifth matching value mentioned above may include:
[0086] If the first tag of the candidate visual media includes a word segment of the search query, or if the second tag of the search query is included, then the fifth matching value corresponding to the candidate visual media is determined to be true.
[0087] If the first tag of the candidate visual media does not include the word segment of the search query, and the second tag of the search query is not included, then the fifth matching value corresponding to the candidate visual media is determined to be false.
[0088] In one possible design approach, the second target matching value may include a sixth matching value, the image information of the candidate visual media includes the similarity between the candidate visual media and the search statement, and the information of the search statement includes the search statement.
[0089] Accordingly, the process of determining the sixth matching value mentioned above may include:
[0090] If the third preset vocabulary includes the search statement, and the similarity between the candidate visual media and the search statement is greater than or equal to the third similarity threshold, then the sixth matching value corresponding to the candidate visual media is determined to be true; the third similarity threshold refers to the similarity threshold corresponding to the search statement in the third preset vocabulary.
[0091] If the similarity between the candidate visual media and the search statement is less than the third similarity threshold, the sixth matching value corresponding to the candidate visual media is determined to be false.
[0092] Optionally, the similarity between the candidate visual media and the search statement used by the electronic device to determine the sixth matching value refers to the similarity between the candidate visual media and the search statement corresponding to the fine-tuning model, and the text feature vector of the search statement refers to the text feature vector of the search statement corresponding to the fine-tuning model.
[0093] In one possible design approach, after obtaining the search query, the electronic device can filter the search query to obtain a filtered search query by filtering out non-visual semantic subjects within the search query. Then, the electronic device can determine the matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information from the filtered search query.
[0094] The first visual medium is determined based on the matching value corresponding to the candidate visual medium. Furthermore, the electronic device can determine a second visual medium that matches the non-visual semantic subject in the search query.
[0095] Then, the electronic device can use the intersection between the first visual media and the second visual media as the target visual media to accurately determine the search results.
[0096] In one possible design approach, the first interface is displayed in response to the operation of opening the gallery application;
[0097] Alternatively, in response to an operation to open the negative one screen triggered by the main screen of the electronic device, the first interface is displayed;
[0098] Alternatively, in response to a pull-down search operation triggered on the main screen of the electronic device, the first interface is displayed.
[0099] In a second aspect, this application provides an electronic device, the electronic device including a display screen, a memory and one or more processors; the display screen, the memory and the processor are coupled; the display screen is used to display an image generated by the processor, the memory is used to store computer program code, the computer program code including computer instructions; when the processor executes the computer instructions, the electronic device performs the method described above.
[0100] Thirdly, this application provides a computer storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform the method described above.
[0101] Fourthly, this application provides a computer program product that, when run on an electronic device, causes the electronic device to perform the method described above.
[0102] It is understood that the beneficial effects achieved by the electronic device described in the second aspect, the computer storage medium described in the third aspect, and the computer program product described in the fourth aspect can be referred to the beneficial effects in the first aspect and any of its possible design embodiments, which will not be repeated here. Attached Figure Description
[0103] Figure 1AA schematic diagram of a visual media search interface provided in an embodiment of this application;
[0104] Figure 1B A second schematic diagram of a visual media search interface provided in an embodiment of this application;
[0105] Figure 1C Schematic diagram three of a visual media search interface provided for an embodiment of this application;
[0106] Figure 1D A schematic diagram of a visual media search interface provided in an embodiment of this application. Figure 4 ;
[0107] Figure 2A A structural block diagram of an electronic device provided in an embodiment of this application;
[0108] Figure 2B A software structure diagram of an electronic device provided in an embodiment of this application;
[0109] Figure 3A Schematic diagram five of a visual media search interface provided for an embodiment of this application;
[0110] Figure 3B A schematic diagram of a visual media search interface provided in an embodiment of this application. Figure 6 ;
[0111] Figure 4 A schematic diagram of a visual media search method provided in an embodiment of this application;
[0112] Figure 5A A schematic diagram illustrating the process of determining an image feature vector as provided in an embodiment of this application;
[0113] Figure 5B Schematic diagram 2 illustrating the process of determining an image feature vector as provided in an embodiment of this application;
[0114] Figure 5C A schematic diagram illustrating the process of determining a text feature vector as provided in an embodiment of this application;
[0115] Figure 5D A schematic diagram illustrating a similarity determination process provided in an embodiment of this application;
[0116] Figure 6 Schematic diagram 2 of a visual media search method provided in an embodiment of this application;
[0117] Figure 7 Schematic diagram three illustrating a visual media search method provided in an embodiment of this application;
[0118] Figure 8A schematic diagram of a visual media search method provided in this application embodiment. Figure 4 ;
[0119] Figure 9A A schematic diagram of a secondary confirmation provided for an embodiment of this application;
[0120] Figure 9B A schematic diagram of a visual media provided for an embodiment of this application;
[0121] Figure 10 Schematic diagram five illustrating a visual media search method provided in an embodiment of this application;
[0122] Figure 11A A schematic diagram of a matrix provided for an embodiment of this application;
[0123] Figure 11B Schematic diagram 2 of a visual media provided in an embodiment of this application;
[0124] Figure 11C A second schematic diagram of a matrix provided for an embodiment of this application;
[0125] Figure 11D A schematic diagram illustrating a positive and negative sample differentiation process provided in an embodiment of this application;
[0126] Figure 11E This is a schematic diagram of a binary classification model training process provided in an embodiment of this application;
[0127] Figure 12 A second schematic diagram illustrating a secondary confirmation provided in this application embodiment;
[0128] Figure 13 Schematic diagram three illustrating a secondary confirmation provided in this application embodiment;
[0129] Figure 14A This application provides an illustration of a secondary confirmation process. Figure 4 ;
[0130] Figure 14B This is a schematic diagram of a secondary confirmation provided in an embodiment of this application. Detailed Implementation
[0131] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "multiple" means two or more.
[0132] To better understand the embodiments of this application, the terminology involved in this application will be explained below.
[0133] Visual media: refers to images or videos.
[0134] Semantic Subjects: Named entity recognition (NER) technology can identify entities with specific meanings in text, such as names of people and places. In this scheme, the identified entities with specific meanings are referred to as semantic subjects.
[0135] Visual content-related and visual content-independent: Visual content refers to the target presented by visual media and the relationships between them. Simply put, visual content can be understood as the content included in visual media. In the context of image search in this application, data that can be obtained from visual media files through the model's natural image understanding is called "visual content-related." This solution refers to data that is related to visual media files but can be obtained without the model's image understanding capabilities as "visual content-independent." For example, electronic devices can acquire and save information such as shooting location, shooting time, name, and file attributes when capturing visual media files.
[0136] For example, in "Photo taken in City 1 this year", "this year" (shooting time), "City 1" (shooting location), and "photo" (file attribute) are all data that electronic devices can acquire and save when capturing visual media files. Therefore, "this year", "City 1", and "photo" are not related to visual semantics. In "Sky taken in City 1 this year", "sky" needs to be understood by the model's image understanding ability to be perceived. Therefore, "sky" is related to visual semantics.
[0137] Text semantic vector: This is a vector that represents the semantic features of the entire sentence, obtained by feeding text into a text encoder. The text encoder can use a clip model or other models, such as the Transformer model commonly used in natural language processing (NLP). This approach does not impose any restrictions on this. In this application, the text semantic vector can also be referred to as the text feature vector.
[0138] Visual semantic vectors: These can be obtained by feeding visual media (such as images) into an image encoder. The image encoder can use a clip model or other models, such as a CNN model or a VIT model; this approach does not impose any restrictions. In this application, visual semantic vectors can also be referred to as image feature vectors.
[0139] Vector similarity: Describes the degree of similarity between two vectors (e.g., between a text semantic vector and a visual semantic vector). In this embodiment, the visual media matching the search query can be determined by comparing the similarity between the text semantic vector of the search query and the visual semantic vector of the visual media. Generally, vector similarity can be calculated using the cosine similarity formula; however, it can also be calculated using other methods.
[0140] Electronic devices (such as mobile phones) can manage users' images, videos, and other visual media through gallery applications. Taking a photo taken with a mobile phone as an example, after the photo is captured, the gallery application determines and stores attribute tags such as the shooting location, shooting time, and photo name, and can use these attribute tags as an index for the image. Once the gallery application has indexed the visual media, it can provide users with corresponding search services. Specifically, users can search for images or videos on their phones by entering keywords in the gallery application. For example, users can enter keywords such as "sky," "cat," or "time point 1" in the search box provided by the gallery application. The gallery application will then match the user's entered keywords with the index of images, videos, and other visual media within the gallery application to obtain search results.
[0141] Optionally, the aforementioned attribute tags may also include attributes such as the person's facial designation (ID), name, and relationship between that person and the mobile phone user. The facial designation of the person in the photo can be automatically generated by the photo library application, while the person's name and relationship can be manually entered by the user. In practice, the same person corresponds to the same facial designation, name, and relationship with the mobile phone user. Therefore, to simplify user operation, the user only needs to enter the person's name and relationship once for each person. Subsequently, the photo library application will automatically configure the person's facial designation, name, and relationship with the mobile phone user for images containing that person's face using facial recognition technology. Additionally, the aforementioned photo name and other attribute information can also be automatically generated by the photo library application or manually named by the user.
[0142] The following is an illustrative description of the interface involved in the search process of the gallery application, with reference to the accompanying drawings:
[0143] like Figure 1A As shown in (a), the mobile phone can display a main interface 101, which can also be referred to as the desktop. The main interface 101 may include an icon 102 for the gallery application. The mobile phone receives a user's click on icon 102; in response to this action, the mobile phone can launch the gallery application and display... Figure 1AThe interface 103 shown in (b) can be a photo album interface. It should be noted that, in response to the user clicking icon 102, the phone can launch the gallery application and display the photo gallery interface. The photo gallery interface includes thumbnails of photos (i.e., images) in the gallery or a large image of a specific photo. Within the photo gallery interface, in response to the user's operation of the "Album" control, the aforementioned album interface 103 is displayed.
[0144] like Figure 1A As shown in (b), interface 103 includes multiple albums, including "All Photos" album containing 2023 photos, "Camera" album containing 1502 photos and videos, "Screenshot and Screen Recording" album containing 102 photos and videos, "My Favorites" album containing 48 photos and videos, "One Record, Multiple Views" album containing 34 photos and videos, "Video Editing" album containing 65 videos, "Custom Albums" album containing 57 photos and videos, and "Shared Albums" album containing 100 photos and videos.
[0145] like Figure 1A As shown in (b), interface 103 may include a search box 104. The mobile phone can receive the user's click on the search box 104, and in response to this action, the mobile phone can display as shown in Figure (b). Figure 1A Interface 105, shown in (c) of the diagram, can be referred to as a search interface. Interface 105 displays photo categorization information to the user. For example, in interface 105, the phone categorizes its photos by time, people, and objects. For instance, in the time dimension, the phone categorizes its photos by three time periods: "This Month," "Last Month," and "This Year." The "This Month" album includes photos or videos taken this month, the "Last Month" album includes photos or videos taken last month, and the "This Year" album includes photos or videos taken this year. In the people dimension, the phone categorizes its photos by different people, such as the four different people shown in interface 105. In the objects dimension, the phone categorizes its photos by "Landscape," "Animal," "Document," and "Architecture." It should be noted that the above categorization dimensions can also be other than those specified here. Users can see this categorization information in interface 105 without entering keywords.
[0146] Optionally, the interface 105 may also include search history 107 and a "clear" option 108. Search history includes keywords previously entered by the user, such as "flowers," "coffee," and "cat." The phone can receive the user's click on "clear" 108, and in response, the phone can clear the search history. After the phone clears the search history, the search interface 105 will no longer display the keywords previously entered by the user. For example, in response to the user's click on "clear" 108, such as... Figure 1B As shown, the search history 107 and "clear" 108 options are no longer displayed on the search interface 105, and the content displayed below has moved up.
[0147] In response to the user entering the keyword "sky" on interface 105, the phone displays the following: Figure 1C Interface 109 is shown in (a). The mobile phone can search for data related to the keyword "sky" on its device. Specifically, the phone can generate related terms such as "sky" and photos containing the word "sky" by associating the keyword "sky". Then, it searches according to each related term, obtaining search results for each related term. For example: 100 photos related to "sky" and 32 photos related to photos containing the word "sky". The 100 photos related to "sky" are retrieved because their respective category tags match "sky" or "sky" and its related terms; the 32 photos related to photos containing the word "sky" are retrieved because optical character recognition (OCR) technology identifies that these 32 photos contain the character "sky". The union of the search results for each of these related terms can be used as the search result for the keyword "sky".
[0148] Interface 109 also displays some search results for the keyword "sky" and "More" options 110 corresponding to the search results for the keyword "sky". The mobile phone receives the user's click on the "More" option 110 and displays the following... Figure 1C Interface 111 is shown in (b) above. Interface 111 displays photos and videos from search results for the keyword "sky". Optionally, photos and videos can be categorized by time. Interface 111 also includes a back button 112 and a title 113. In response to the user's operation of the back button 112, the phone can redisplay interface 109. The title 113 may include the keyword "sky".
[0149] That is to say, in the gallery application, when a user enters a simple search statement in the search box, such as a simple keyword, for example: sky, location 1, time 1, etc., corresponding search results can be obtained. However, because the mobile phone's ability to understand and associate search statements is limited, when the user enters a relatively complex search statement in the search box, if the keywords in the search statement cannot match the attribute tags of the pictures or the text in the pictures, no photos may be found. That is to say, the existing gallery applications do not support the search function based on complex search statements. As Figure 1D shown, when the user enters a relatively complex search statement "warming oneself by the fire and brewing tea" in the search box of interface 114, the mobile phone cannot understand the associated words of "warming oneself by the fire and brewing tea", and since there are no tags on the photos that match "warming oneself by the fire and brewing tea" or its associated words, the search results show "no pictures".
[0150] In some embodiments, the electronic device can use an encoder to calculate the similarity between the visual media on the electronic device and the search statement, so that the electronic device can use the visual media with a similarity higher than the threshold as the search result. Then, the electronic device displays the search result to present the visual media required by the user. For example, the encoder can include a text encoder and an image encoder. The electronic device can input each visual media on the electronic device into the image encoder to obtain the image feature vector corresponding to each visual media. And the electronic device can input the search statement into the text encoder to obtain the text feature adjacent corresponding to the search statement. Then, for each visual media, the electronic device can calculate the similarity between the visual media and the search statement based on the image feature vector corresponding to the visual media and the text feature vector corresponding to the search statement.
[0151] For example, the encoder described above could be a single encoder. An electronic device can input various visual media and search terms into this encoder, allowing it to calculate the similarity between each visual media and the search term based on the image feature vectors of the visual media and the text feature vectors of the search term. However, relying solely on the encoder may not accurately determine the image feature vectors of the visual media, resulting in low accuracy in calculating the similarity between the visual media and the search term. Consequently, the search results displayed by the electronic device may not be the visual media the user needs, failing to meet the user's search requirements. For instance, when the search term includes landmarks, the encoder may not accurately determine the image feature vectors of visual media containing landmarks, leading to low accuracy in calculating the similarity between the visual media and the search term, and consequently, the electronic device cannot accurately search for the visual media the user needs. Simply put, the electronic device cannot accurately identify visual media containing landmarks, resulting in low accuracy in the determined search results. Similarly, when the search term includes uncommon animals, the encoder may not accurately identify visual media containing animals not present in the scene, leading to low accuracy in the determined search results.
[0152] Therefore, to improve the accuracy of visual media search and meet users' search needs, this application provides a visual media search method. After identifying visual media with similarity higher than a threshold using a relevance model (such as a multimodal model), the electronic device uses these visual media as candidate visual media. Subsequently, the electronic device can further verify the candidate visual media to determine whether they are the visual media required by the user, thus obtaining the target visual media and achieving accurate determination of the search results. Finally, the electronic device can display the target visual media, i.e., display the visual media required by the user, satisfying the user's search needs.
[0153] For example, the aforementioned electronic devices can be mobile phones, tablets, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), and other devices capable of storing visual media.
[0154] For example, Figure 2AA schematic diagram of the structure of electronic device 200 is shown. Electronic device 100 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, antenna 1, antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone jack 270D, a sensor module 280, buttons 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.
[0155] The sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an accelerometer sensor 280E, a distance sensor 280F, a proximity sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.
[0156] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 200. In other embodiments of this application, the electronic device 200 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0157] Processor 210 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.
[0158] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.
[0159] The processor 210 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory can store instructions or data that the processor 210 has just used or that are used repeatedly. If the processor 210 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.
[0160] In some embodiments, the processor 210 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0161] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the electronic device 200. In other embodiments of this application, the electronic device 200 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0162] Electronic device 200 implements display functions through a GPU (Graphics Processing Unit), a display screen 294, and an application processor. The GPU connects the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 210 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0163] Display screen 294 is used to display images, videos, etc. Display screen 294 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 294, where N is a positive integer greater than 1.
[0164] Electronic device 200 can perform shooting functions through ISP, camera 293, video codec, GPU, display screen 294 and application processor.
[0165] Camera 293 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 293, where N is a positive integer greater than 1.
[0166] Video codecs are used to compress or decompress digital video. Electronic device 200 may support one or more video codecs. Thus, electronic device 200 can play or record video in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0167] An NPU (Neural Processing Unit) is a neural network (NN) computing processor that, by borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, rapidly processes input information and can continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0168] The external storage interface 220 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 200. The external memory card communicates with the processor 210 through the external storage interface 220 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.
[0169] Internal memory 221 can be used to store computer executable program code, which includes instructions. Internal memory 221 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 200 (such as audio data, phonebook, etc.). Furthermore, internal memory 221 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 210 executes various functional applications and data processing of electronic device 200 by running instructions stored in internal memory 221 and / or instructions stored in memory disposed in the processor.
[0170] Figure 2B This is a software structure block diagram of an electronic device 200 according to an embodiment of this application. The software system of the electronic device 200 can adopt a layered architecture, which divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. Taking the Android system as an example, in some embodiments, the Android system can be divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0171] like Figure 2B As shown, the application layer may include applications such as a gallery service module, a search module, a multimodal understanding module, a natural language understanding module, and a camera application.
[0172] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0173] The system library can include multiple functional modules, such as a surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), and 2D graphics engines (e.g., SGL). The media libraries support playback and recording of various common audio and video formats, as well as still image files. The media libraries support multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0174] The kernel layer is the layer between hardware and software.
[0175] The following example, using a scene of capturing a photograph, illustrates the workflow of the software and hardware of the electronic device 200.
[0176] When the touch sensor 280K receives a touch operation, a corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including touch coordinates, timestamp of the touch operation, etc.). The raw input event is stored in the kernel layer. The application framework layer retrieves the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking a touch click as an example, where the click corresponds to the camera application icon, the camera application calls the interface of the application framework layer to launch the camera application, and then calls the kernel layer to launch the camera driver, capturing still images or videos through camera 293.
[0177] Taking a mobile phone as an example, the visual media search method provided in this application embodiment will be introduced. The visual media search method provided in this application embodiment can be applied to applications such as gallery applications and file management applications.
[0178] The interface and search logic involved in the visual media search method provided in the embodiments of this application will be described exemplarily below with reference to the accompanying drawings.
[0179] like Figure 3A Interface 301, shown in (a), displays search history 303 and a "clear" option 304. Search history 303 includes previously entered search terms, such as "sunrise from the mountaintop" and "sky photographs." Other content displayed on interface 301 can be referenced from the related content on interface 105, and will not be repeated here. Interface 301 can be referred to as the search interface. The mobile phone can respond to user queries... Figure 1A Clicking on the search box 104 in the interface 103 shown in (b) in
[0180] As Figure 3A shown in (b) in Figure 3A the interface 305 (i.e., the search result interface), the user enters the search statement "making tea around the stove" in the search box 306 of the interface 305. The mobile phone searches for 239 pictures. The mobile phone displays some search results of the search statement "making tea around the stove" (e.g., thumbnails of 8 pictures) and the "More" option 307 corresponding to the search results of the search statement "making tea around the stove" on the interface 305. In response to the user's operation on the "More" option 307, the mobile phone displays the interface 308 shown in (c) in
[0181] In addition, the mobile phone can also be provided with a negative first screen, a pull-down search interface, etc. It can be understood that the negative first screen can be the leftmost split screen of the electronic device, used to provide functions such as search and quick services for users. Among them, the negative first screen can also be used to display notification messages to be pushed to users, such as application messages subscribed by users, real-time hot search messages, segment selection, itinerary information, etc. The pull-down search interface is an interface displayed in response to the user's pull-down operation on the main interface. This interface is used to provide functions such as search and application suggestions for users. This interface and the Figure 3B interface 315 in (c) below can be the same interface.
[0182] Next, an example will be given with the negative first screen. When the user needs to view the negative first screen of the mobile phone, the user can slide the screen of the mobile phone to make the electronic device display the negative first screen.
[0183] Exemplarily, as shown in (a) in Figure 3B the mobile phone can receive operation 1 implemented by the user on the interface 309 (which can be called the desktop) of the mobile phone. Exemplarily, this operation 1 can be a rightward sliding operation as shown in (a) in Figure 3B In response to this operation 1, the mobile phone can display the negative first screen 310 shown in (b) in Figure 3B Among them, the negative first screen 310 can include: a search box 311, quick services 312, default cards 313, recommended cards 314, etc. The quick services 312 can be quick access to a certain page or function of an application program, such as: scan code, payment code, ride code, etc.; the default cards can be: gallery cards, remaining battery cards, etc.; the recommended cards can be recommended application cards.
[0184] The mobile phone receives a click from the user on the search box 311 on the negative one screen 310. In response to this click, the mobile phone can display as follows: Figure 3B Interface 315 is shown in (c). Interface 315 may include: a search box 316 and application suggestions. Application suggestions include icons of suggested applications. Interface 315 may also include: search history 317 and its corresponding "clear" option 318. In response to the user's triggering operation on the "clear" option 318, the search history 317 and the "clear" option 318 will no longer be displayed on interface 315. In addition, the search box 316 may display trending news headlines, such as "marathon race".
[0185] like Figure 3B As shown in (d) of the diagram, interface 319 displays the user-inputted search term "mountain peaks photographed on weekends" in its search box. Interface 319 also displays a preview area 322 of the search results for "mountain peaks photographed on weekends" from the photo library app, as well as a "Search in app" option 323 from the photo library app. In response to the user's action on preview area 322, the phone can access the photo details interface provided by the photo library app, allowing the user to browse and view the search results for "mountain peaks photographed on weekends," including visual media such as photos or videos matching the search term. In response to the user's action on "Search in app" option 323, the phone displays... Figure 3B As shown in (e), interface 324, provided by the photo library application, displays partial search results for the search query "mountains photographed on weekends" and a "more" option corresponding to the search results for the same query. In response to a user's triggering of this "more" option, the phone can display a search results details interface, showing images from the search results for "mountains photographed on weekends." Additionally, interface 319 may also display an online search option 321. In response to a user's triggering of the online search option 321, the phone displays a search webpage and shows the online search results within that webpage.
[0186] The above section introduced the interface and search logic involved in the visual media search method. The following section will continue to combine the above... Figure 2B The software structure shown illustrates the specific implementation process of the visual media search method. For example... Figure 4 As shown, the implementation process may include S401-S423, where S401-S407 may belong to the index building stage, and S408-S423 may belong to the search stage.
[0187] S401. Add and / or modify visual media and their attributes in the Gallery Service module.
[0188] The attributes of visual media may include one or more of the following: capture location, capture time, visual media name, facial ID of a person in the visual media, person's name, and the relationship between that person and the mobile phone user. For example, in the case of a captured video or image, the capture location refers to the location where the video was taken, and the capture time refers to the time the video was taken; in the case of a screenshot, the capture location refers to the location where the screenshot was taken, and the capture time refers to the time the screenshot was taken; in the case of a downloaded video or image, the capture location refers to the location where the video was downloaded, and the capture time refers to the time the video was downloaded.
[0189] Users can add new visual media by taking photos, downloading, or taking screenshots. In addition, users can modify existing visual media to create new ones. These modifications include, but are not limited to, enhancements, custom naming, and adding watermarks.
[0190] S402, The image library service module stores visual media and their attributes.
[0191] The gallery service module can respond to user input regarding adding or modifying visual media, storing the visual media and its attributes locally on the phone. In practical applications, with user authorization, the phone can store the locally stored visual media and its attributes in the cloud to reduce the storage pressure on the phone's local storage.
[0192] S403, the image library service module sends request 1 to the multimodal understanding module. Request 1 is used to trigger visual semantic understanding of visual media.
[0193] S404. The multimodal understanding module returns the image feature vector of the visual media to the image library service module.
[0194] In this embodiment, the multimodal understanding module responds to request 1 above and determines the image feature vector of the visual media stored on the mobile phone. Since visual semantic understanding requires significant computing resources, to avoid impacting user experience, the gallery service module can request the multimodal understanding module to perform visual semantic understanding on the visual media stored locally on the mobile phone while the phone is charging and the screen is off. This includes performing visual semantic understanding on newly added or modified visual media to obtain the visual semantic vector (or image feature vector) of the visual media. This enables offline processing of the visual media, reducing the impact of visual semantic understanding on other services running on the mobile phone.
[0195] The multimodal understanding module can perform visual semantic understanding of visual media based on a multimodal model, obtaining image feature vectors of the visual media. Furthermore, the multimodal model can not only be used for visual semantic understanding of visual media to obtain visual semantic vectors, but also for semantic understanding of search statements to obtain text feature vectors (or sentence semantic vectors) of the search statements. The process of semantic understanding of search statements by the multimodal model can be referred to in the relevant description below, and will not be detailed here.
[0196] In some embodiments, the multimodal model can map visual media and text into vectors of the same dimension; that is, the dimension of the visual semantic vector of the visual media is the same as the dimension of the semantic vector of the text (e.g., the sentence semantic vector of a search query). Specifically, the multimodal model can be based on a contrastive language-image pre-training (CLIP) model. Mobile phones can use the CLIP model to map visual media and text into a unified vector space to understand the relationships between different modalities in both text and vision, thereby enabling image retrieval.
[0197] Among them, the CLIP model is either a standard CLIP model, that is, an existing CLIP model, or a custom CLIP model.
[0198] For example, the aforementioned custom CLIP model may include a base model and a fine-tuned model. The fine-tuned model is obtained by further training the base model using a training set for a specific scenario. Therefore, the fine-tuned model has a limited scope, capable of processing images and text in a specific scenario. The generalization ability of the fine-tuned model is less than that of the base model, while the retrieval ability of the base model in a specific scenario is less than that of the fine-tuned model.
[0199] In some embodiments, the base model and fine-tuning model reuse parts of the network to reduce resource consumption on mobile phones with a custom CLIP model installed. The base model and fine-tuning model need to process not only visual media but also text. The module in the base model that processes visual media can be called the visual encoding module, and the module in the fine-tuning model that processes visual media can be called the visual fine-tuning module. The image feature vectors mentioned above can include the image feature vectors corresponding to the base model and the image feature vectors corresponding to the fine-tuning model. The visual encoding module of the base model and the visual fine-tuning module of the fine-tuning model can reuse the image encoder, allowing the visual encoding module to use the feature vectors of the visual media output by the image encoder to further determine the first L2 norm of the visual media, thus determining the image feature vectors corresponding to the base model. The visual fine-tuning module can use the feature vectors of the visual media output by the image encoder to further determine the second L2 norm of the visual media, thus determining the image feature vectors corresponding to the fine-tuning model.
[0200] For example, upon receiving request 1 from the gallery service module, the response to request 1 is as follows: Figure 5A As shown, the multimodal understanding module can input k visual media from the mobile phone into the visual encoding module. The image encoder in the visual encoding module encodes each of the k visual media to obtain the image feature vector 1 of each visual media. This image feature vector can be 768-dimensional, i.e., X = {x1, x2, ..., x768}, where X represents the image feature vector 1. Then, on one hand, the visual encoding module can output this 768-dimensional image feature vector 1 so that it can continue to be used as the input parameter of the visual fine-tuning module. On the other hand, the visual encoding module continues to use the mapping matrix 1 and the image feature vector 1 to calculate and output the first L2 norm α1 of the image feature vectors of each of the k visual media. The mapping matrix 1 has a dimension of 768*512 to map the image feature vector 1 from 768 dimensions to 512 dimensions. Here, the image feature vector 1 and the first L2 norm of the visual media can be considered as the image feature vectors of the visual media corresponding to the base model.
[0201] Specifically, for each of the k visual media, the visual encoding module can adopt... Calculate the first L2 norm α1 of the image feature vectors of visual media. Here, X represents a 768-dimensional image feature vector 1, M1 represents a mapping matrix 1, and V1 represents an image feature vector 2, where V1 has a dimension of 512. This represents the i-th element in V1.
[0202] like Figure 5BAs shown, after receiving the image feature vector 1 of each of the k visual media, the visual fine-tuning module calculates and outputs the second L2 norm α2 of the image feature vector for each visual media using the mapping matrix 2 and the image feature vector 1. The mapping matrix 2 has a dimension of 768*512 to map the image feature vector 1 from 768 dimensions to 512 dimensions. Here, the image feature vector 1 and the second L2 norm of the visual media can be considered as the image feature vector of the visual media corresponding to the fine-tuning model.
[0203] Specifically, for each of the k visual media, the visual encoding module can adopt... Calculate the second L2 norm α2 of the image feature vectors of visual media. Here, X represents the 768-dimensional image feature vector 1, M2 represents the mapping matrix 2, and V2 represents the image feature vector 3, where V2 has a dimension of 512. This represents the i-th element in V2.
[0204] In some embodiments, the fine-tuning model described above (such as mapping matrix 2 in the visual fine-tuning module) is obtained by further training using a training set for a specific scene. When training the fine-tuning model, the device can freeze the image encoder and train only the last fully connected layer M2 of the fine-tuning model. The training set for this specific scene can be an image training set including visual content corresponding to a preset whitelist. This preset whitelist represents the support range of the fine-tuning model. The device can be the aforementioned mobile phone, or it can be something else; this application does not limit the device used to train the fine-tuning model.
[0205] After obtaining the image feature vectors of the visual media (such as the image feature vectors of the visual media corresponding to the base model and the image feature vectors of the visual media corresponding to the fine-tuning model), in order to facilitate the retrieval of the visual media required by the user using the image feature vectors of the visual media, the multimedia understanding module can save the image feature vectors of the visual media, namely the image feature vector 1, the first L2 norm α1, and the second L2 norm α2 mentioned above. This allows the storage of an image feature vector of a visual media to only require saving a 770-dimensional visual vector, instead of saving a 512*2-dimensional, i.e., 1024-dimensional visual feature vector, thus reducing the resources required to store the image feature vectors. Among them, the 512-dimensional vector is the image feature vector obtained by processing the 768-dimensional vector using mapping matrix 1 or mapping matrix 2 (i.e., the V1 and V2 mentioned above).
[0206] It should be noted that the image encoder described above, located in the visual encoding module, is merely an example. The image encoder could also be located in the visual fine-tuning module, and this application does not limit it. Furthermore, the image feature vector of the visual media described above, including the image feature vector of the visual media corresponding to the base model and the image feature vector of the visual media corresponding to the fine-tuning model, is only one example. The image feature vector of the visual media could also include only one image feature vector determined by a single model, which could be a clip model or not.
[0207] As described above, the image feature vectors of visual media on a mobile phone can be determined offline, while the text feature vectors of the search query can be determined online after the mobile phone receives the user's search query. This allows the mobile phone to use the text feature vectors of the search query and the image feature vectors of the visual media on the phone to determine the visual media that matches the search query. The process of determining the text feature vectors of the search query and using them to determine the matching visual media is detailed below and will not be elaborated on here. We will now continue by introducing the process of building an index for visual media.
[0208] S405, the image library service module stores the image feature vectors of visual media.
[0209] S406. The image library service module sends the attribute information of the visual media and its visual semantic vector to the search module.
[0210] In this embodiment, the image library service module can locally store the image feature vector of each of the k visual media on the mobile phone. Furthermore, the image library service module can send the attribute information and image feature vectors of the visual media in batches to the search module, so that the search module can construct an index of the visual media.
[0211] Optionally, the multi-image library service module may choose not to upload the image feature vector of the visual media to the cloud. Of course, it may also upload the image feature vector of the visual media to the cloud with the user's authorization. This application does not restrict this.
[0212] S407, The search module builds an index for visual media.
[0213] The index of visual media built by the search module may include: the attributes of the visual media and / or the visual semantic vector of the visual media.
[0214] S408, The image library service module receives the search query entered by the user.
[0215] Users can use the search interface provided by the gallery service module, such as the search results mentioned above. Figure 3AEnter "warming the tea around the stove" in the search box 306 on the interface 305 shown in (b) in []. This "warming the tea around the stove" is the search statement.
[0216] S409. The gallery service module sends the search statement to the search module.
[0217] S410. The search module determines whether the search statement includes visual content.
[0218] In the embodiment of the present application, the search module can determine whether the search statement includes visual content, that is, determine whether it is necessary to search for visual media using image feature vectors and text feature vectors. In other words, the search module determines whether it is necessary to use a multimodal model to determine visual media.
[0219] When the search statement does not include visual content, it indicates that visual media can be searched using the attributes of visual media, without the need to search for visual media using image feature vectors and text feature vectors, that is, it indicates that there is no need to use a multimodal model to determine visual media, and the search module can execute S411.
[0220] When the search statement includes visual content, it indicates that it is necessary to search for visual media using image feature vectors and text feature vectors, that is, it indicates that it is necessary to use a multimodal model to determine visual media, and the search module can execute S412.
[0221] In some embodiments, the search module can determine whether a search statement includes visual content by judging whether it includes a visual semantic subject. The search module can send request 2 to the natural language understanding module. The natural language understanding module can perform semantic subject recognition on the search statement to obtain the semantic subjects included in the search statement. The natural language understanding module performs semantic subject recognition based on a natural language understanding model. Specifically, the natural language understanding module can use named entity recognition (NER) technology to perform semantic subject recognition on the search statement to obtain the semantic subjects contained in the search statement. In this embodiment, a semantic subject can also be referred to as an entity, and a semantic subject may include one or more of the following: semantic subjects related to time, semantic subjects related to location, semantic subjects related to names, semantic subjects related to relationships between people, and semantic subjects related to visual content (or visual semantic subjects). Optionally, the semantic subjects related to visual content are determined from a preset set of M (M≥1) semantic subjects related to visual content using named entity recognition technology. These M semantic subjects can be configured by the developers of the image library service module according to actual needs. Generally, these M semantic subjects are all nouns. For example, semantic subject identification of the search query "sky taken in city 1 on September 1st" can yield three semantic subjects: "September 1st", "city 1", and "sky". Among them, "sky" is related to visual content and can be a visual semantic subject.
[0222] Next, the natural language understanding module can return the identified semantic subject to the search module. The search module then determines whether the semantic subject includes a visual semantic subject. If the semantic subject does not include a visual semantic subject, the search module can determine that the search query does not include visual content. If the semantic subject includes a visual semantic subject, the search module can determine that the search query includes visual content.
[0223] S411. The search module retrieves search results based on the index and search query of visual media.
[0224] For example, if the search query does not include a visual semantic subject, indicating that the search query does not include content related to visual semantics, the search module can directly query the visual media corresponding to the semantic subject based on the constructed index, that is, based on the attributes of each visual media, and use it as the search result. Here, the semantic subject refers to the attribute of the visual media. For example, if the search query is "September 1st" and does not include visual content, the search module can use the index to find visual media whose collection time is September 1st and obtain the search result. As another example, if the search query includes a photo of Zhang San but does not include visual content, the search module can match the names of people in each visual media with "Zhang San" to determine the visual media whose name attribute is "Zhang San" and use it as the search result.
[0225] S412. The search module filters the non-visual semantic subjects in the above search statement to obtain the filtered search statement.
[0226] For example, when the aforementioned semantic subjects include visual semantic subjects, it indicates that the search statement includes content related to visual semantics and may also include content unnecessary for visual semantics, i.e., non-visual semantic subjects. Since non-visual semantic subjects are irrelevant to the search for visual content, the search module can first filter the non-visual semantic subjects in the search statement to obtain a filtered search statement, in order to search for visual media that match the filtered search statement.
[0227] Optionally, non-visual semantic subjects refer to semantic subjects related to attributes of visual media, such as semantic subjects related to time, semantic subjects related to location, semantic subjects related to names, and semantic subjects related to relationships between people. It should be understood that since relationships between people have already been considered as attributes of visual media, they can be considered non-semantic subjects here.
[0228] For example, in the search query "sky photographed this year", "this year" is a non-visual semantic subject, so the filtered search query is "sky photographed".
[0229] In practical applications, after filtering semantic subjects unrelated to visual content, some redundant stop words may remain. For example, in the search query "sky photographed in city 1 this year," after deleting "this year" and "city 1," the stop word "in" becomes redundant and can therefore be filtered out by the search module. Specifically, the search module can filter semantic subjects unrelated to visual content and their associated stop words in the search query to obtain a filtered search query. For example, the filtered search query corresponding to the search query "sky photographed in city 1 this year" is "the sky photographed."
[0230] S413, The search module sends a filtered search statement to the multimodal understanding module.
[0231] S414 The multimodal understanding module performs semantic understanding on the filtered search statement to obtain the text feature vector of the filtered search statement.
[0232] For example, the multimodal understanding module can employ a multimodal model to determine the text feature vectors of the filtered search statements. Optionally, the multimodal model can be a CLIP model. The CLIP model can be a standard CLIP model, or the CLIP module can be a custom CLIP model.
[0233] In some embodiments, as described above, the image feature vector may include the image feature vector corresponding to the base model and the image feature vector corresponding to the fine-tuning model. The visual encoding module of the base model and the visual fine-tuning module of the fine-tuning model can reuse the image encoder. Accordingly, in order to maintain the similarity of the image-text pairs, this application introduces a text encoding module, which can reuse the text encoder to output the text feature vector of the filtered search statement corresponding to the base model and the text feature vector of the filtered search statement corresponding to the fine-tuning model, respectively. This allows for the calculation of the similarity between the search statement and the visual media using the text feature vector of the filtered search statement corresponding to the base model and the image feature vector of the visual media corresponding to the base model, and the calculation of the similarity between the search statement and the visual media using the text feature vector of the filtered search statement corresponding to the fine-tuning model and the image feature vector of the visual media corresponding to the fine-tuning model, thereby realizing the calculation of the similarity of the image pairs.
[0234] For example, such as Figure 5C As shown, the multimodal understanding model inputs the filtered search statement into the CLIP model. The text encoder in the CLIP model's text encoding module encodes the filtered search statement, obtaining and outputting its feature vector, which has 768 dimensions. Subsequently, the text encoding module can input the feature vector into the base model branch and the fine-tuning model branch, respectively. The base model branch and the fine-tuning model branch have similar structures.
[0235] For the base model branch: the text encoding module uses mapping matrix 3 and the feature vector of the filtered search statement to obtain a 512-dimensional intermediate variable 1. For example, the text encoding module calculates intermediate variable 1 according to T1 = YN1. Here, T1 represents intermediate variable 1, Y represents the feature vector of the filtered search statement, N1 represents mapping matrix 3, and the text encoding module uses intermediate variable 1 to calculate the first L2 norm of the filtered search statement. For example, the text encoding module can... Calculate the first L2 norm of the filtered search statement. Where β1 represents the first L2 norm of the filtered search statement. Let i represent the i-th element in T1.
[0236] Since the image feature vector of visual media has a dimension of 768, and the intermediate variable 1 of the filtering search statement has a dimension of 512, the text encoding module needs to map the intermediate variable 1 of the filtering search statement to a 768-dimensional form in order to perform calculations between the image feature vector and the text feature vector. Specifically, the text encoding module can use the aforementioned mapping matrix 1 and the first L2 norm of the filtering search statement to map the intermediate variable 1 into a 768-dimensional text feature vector, thereby obtaining the text feature vector of the filtering search statement corresponding to the base model. Calculate the text feature vector of the filtered search statement corresponding to the base model. Here, T'1 represents the text feature vector of the filtered search statement corresponding to the base model. It is the transpose of M1, where M1 represents the above mapping matrix 1.
[0237] For the fine-tuning model branch: the text encoding module uses mapping matrix 4 and the feature vectors of the filtered search statements to obtain a 512-dimensional intermediate variable 2. For example, the text encoding module calculates intermediate variable 2 based on T2 = YN2. Here, T2 represents intermediate variable 2, N2 represents mapping matrix 4, and the text encoding module uses intermediate variable 2 to calculate the second L2 norm of the filtered search statements. For example, the text encoding module can... Calculate the second L2 norm of the filtered search statement. Where β2 represents the second L2 norm of the filtered search statement. Let i represent the i-th element in T2.
[0238] Since the image feature vector of visual media has a dimension of 768, and the intermediate variable 2 of the filtering search statement has a dimension of 512, the text encoding module needs to map the intermediate variable 2 of the filtering search statement to a 768-dimensional form in order to perform calculations between the image feature vector and the text feature vector. Specifically, the text encoding module can use the aforementioned mapping matrix 2 and the first L2 norm of the filtering search statement to map the intermediate variable 2 into a 768-dimensional text feature vector, thereby obtaining the text feature vector of the filtering search statement corresponding to the fine-tuning model. Specifically, the text encoding module can utilize... Calculate the text feature vector of the filtered search statement corresponding to the base model. Here, T'2 represents the text feature vector of the filtered search statement corresponding to the fine-tuned model. M1 is the transpose of M2, and M1 represents the above mapping matrix 2.
[0239] It should be noted that the text feature vectors of the filtered search statements for the fine-tuning model and the filtered search statements for the base model can be output simultaneously by the text encoding module.
[0240] In this embodiment, the base model and fine-tuning model in the custom CLIP model reuse the text encoder, so that the custom CLIP model only needs to calculate the feature vector of the filtered search statement once. Then, it can use the feature vector of the filtered search statement to determine the text feature vector of the filtered search statement corresponding to the base model and the text feature vector of the filtered search statement corresponding to the fine-tuning model, respectively, without the base model and the fine-tuning model calculating the feature vector of the filtered search statement separately, thereby improving the computational efficiency of the text feature vector of the filtered search statement.
[0241] In this embodiment, to ensure the generalization ability and the ability to recognize specific images and text of the multimodal model, the multimodal model includes a base model and a fine-tuning model. Since the base model and the fine-tuning model have the same model weights, meaning some parts of the network are identical, directly setting two models would waste mobile phone memory. Therefore, this application reuses the image encoder and text encoder for both the base model and the fine-tuning model, allowing them to determine their respective text feature vectors and image feature vectors using the feature vectors output by the image encoder and text encoder, respectively. Overall, this reduces the determination time of image feature vectors and text feature vectors by nearly half.
[0242] S415. The multimodal understanding module performs recall based on text feature vectors and image feature vectors of visual media to obtain candidate visual media.
[0243] In some embodiments, the image feature vector of the aforementioned visual media may be sent by the search module to the multimodal understanding module, or it may be obtained by the multimodal understanding module from the local mobile phone.
[0244] In this embodiment, for each visual medium on the mobile phone (such as each of the k visual media mentioned above), the multimodal understanding module can calculate the vector similarity (or simply similarity) between the image feature vector of the visual medium and the text feature vector of the filtering search statement. That is, based on the image feature vector of the visual medium and the text feature vector of the filtering search statement, the similarity between the visual medium and the filtering search statement is calculated. Then, based on the vector similarity, the multimodal understanding module determines the visual medium matching the filtering search statement from the k visual media and uses the determined visual medium as candidate visual media.
[0245] For example, the multimodal understanding module can identify visual media with a vector similarity greater than or equal to a threshold of 1 as visual media matching the filtering search statement. Optionally, if the number of visual media with a vector similarity greater than or equal to the threshold of 1 is greater than one, the multimodal understanding module can sort the visual media with a vector similarity greater than or equal to the threshold of 1 in descending order of vector similarity. Then, the multimodal understanding module can identify the top n sorted visual media as visual media matching the filtering search statement. Here, n is a positive integer.
[0246] In some embodiments, the image feature vector of the visual media may include the image feature vector of the visual media corresponding to the base model and the image feature vector of the visual media corresponding to the fine-tuning model. The text feature vector of the filtered search statement includes the text feature vector of the filtered search statement corresponding to the base model and the text feature vector of the filtered search statement corresponding to the fine-tuning model. In one case, such as... Figure 5D As shown, the multimodal understanding module can calculate the vector similarity 1 between the text feature vector of the filter search statement corresponding to the base model and the image feature vector of the visual media corresponding to the base model (that is, calculate the similarity between the visual media corresponding to the base model and the filter search statement), and calculate the vector similarity 2 between the text feature vector of the filter search statement corresponding to the fine-tuning model and the image feature vector of the visual media corresponding to the fine-tuning model (that is, calculate the similarity between the visual media corresponding to the fine-tuning model and the filter search statement). Then, the multimodal understanding module can determine visual media with vector similarity 1 or vector similarity 2 greater than or equal to threshold 1 as visual media matching the filter search statement.
[0247] In another scenario, the multimodal understanding module can determine whether the filtered search statement falls within the support range of the fine-tuning model by checking if it hits whitelist 1, thus determining whether to use the fine-tuning model branch to identify candidate visual media. The following will combine... Figure 6 This paper introduces the process by which the multimodal understanding module uses whitelist 1 to determine candidate visual media.
[0248] S501, The multimodal understanding module determines whether the filtered search statement belongs to the whitelist.
[0249] In this embodiment of the application, if the filtered search statement is not in the whitelist, it indicates that the filtered search statement is within the support range of the base model. The multimodal understanding module can use the base model branch to determine the candidate visual media, and the multimodal understanding module can execute S502.
[0250] If the filtered search query is on the whitelist, it indicates that the search query is within the support range of the fine-tuning model. The multimodal understanding module can then determine the candidate visual media by making the fine-tuning model branch, and the multimodal understanding module can execute S504.
[0251] In some embodiments, the multimodal understanding module determines whether each semantic subject (i.e., visual semantic subject) in the filtered search statement belongs to whitelist 1. Considering that user-input search statements are generally phrases, the multimodal understanding module can determine that the filtered search statement belongs to the whitelist if all semantic subjects in the filtered search statement belong to whitelist 1. If any semantic subject does not belong to whitelist 1, the multimodal understanding module determines that the filtered search statement does not belong to the whitelist. For example, if the filtered search statement is "a boy holding a basket," the semantic subjects include "basket" and "boy." The multimodal understanding module can determine whether the basket and the boy belong to whitelist 1 respectively. If both "blue" and "boy" belong to whitelist 1, the multimodal understanding module can determine that the filtered search statement belongs to the whitelist. If neither "blue" nor "boy" belongs to whitelist 1, the multimodal understanding module can determine that the filtered search statement does not belong to the whitelist.
[0252] S502. For each visual medium, the multimodal understanding module calculates the vector similarity 1 between the image feature vector of the visual medium corresponding to the base model and the text feature vector of the filter search statement corresponding to the base model.
[0253] In the embodiments of this application, such as Figure 7 As shown, when the filtered search statement does not belong to the whitelist 1, the multimodal understanding module can call the text feature vector of the filtered search statement output by the base model branch of the text encoding module and the image feature vector of the visual media corresponding to the base model to calculate the score of the visual media corresponding to the base model, that is, calculate the vector similarity 1 between the visual media corresponding to the base model and the filtered search statement.
[0254] Specifically, the multimodal understanding module can be achieved through... Calculate the score corresponding to the visual media of the base model. Here, S1 represents the score corresponding to the visual media of the base model, α1 represents the first L2 norm of the image feature vector of the aforementioned visual media, X represents the image feature vector 1, and T'1 represents the text feature vector of the filtered search statement corresponding to the aforementioned base model.
[0255] S503, the multimodal understanding module selects visual media with a vector similarity of 1 greater than the threshold 1 as candidate visual media.
[0256] S504. For each visual medium, the multimodal understanding module calculates the vector similarity 2 between the image feature vector of the visual medium corresponding to the fine-tuning model and the text feature vector of the filter search statement corresponding to the fine-tuning model.
[0257] In the embodiments of this application, as described above Figure 7 As shown, when the filtered search statement belongs to whitelist 1, the multimodal understanding module can call the text feature vector of the filtered search statement output by the fine-tuning model branch of the text encoding module and the image feature vector of the visual media corresponding to the fine-tuning model to calculate the score corresponding to the visual media corresponding to the fine-tuning model, that is, to calculate the vector similarity 2 between the visual media corresponding to the fine-tuning model and the filtered search statement.
[0258] Specifically, the multimodal understanding module can be achieved through... Calculate the score corresponding to the visual media of the base model. Here, S2 represents the score corresponding to the visual media of the fine-tuned model, and α2 represents the second L2 norm of the image feature vector of the aforementioned visual media. X represents the aforementioned image feature vector 1, and T'2 represents the text feature vector of the filtered search statement corresponding to the aforementioned fine-tuned model.
[0259] S505, the multimodal understanding module selects visual media with a vector similarity of 2 greater than the threshold of 1 as candidate visual media.
[0260] In some embodiments, the operations performed by the multimodal understanding module can be performed by the multimodal model. Additionally, S501-S505 can be performed by the multimodal understanding module, i.e., the model output module within the multimodal model.
[0261] In this embodiment, the multimodal understanding module can determine the image feature vector of visual media offline, while only needing to determine the text feature vector of the search query online. This improves the computation time for vector similarity between image and text feature vectors, effectively reducing user retrieval time. Furthermore, the multimodal understanding module can determine whether to use a fine-tuning model to identify candidate visual media based on whether the filtered search query belongs to whitelist 1, thereby improving the accuracy of candidate visual media identification and ultimately enhancing the accuracy of search results.
[0262] Optionally, the whitelist 1 mentioned above is expandable, meaning the scope of support for the fine-tuning model is expandable. In other words, the scenarios supported by the fine-tuning model are expandable; the fine-tuning model can be trained using the training set corresponding to the expanded scenarios. It should be understood that training the fine-tuning model actually involves training the fine-tuning model branches in the fine-tuning module and the text encoding module, without needing to train the reused parts (such as the image encoder and text encoder mentioned above).
[0263] In some embodiments, after determining candidate visual media, the multimodal understanding module can directly use these candidate visual media as search results. The multimodal understanding module can then return the search results to the search module. The search module can then send the search results to the image library service module for display.
[0264] Furthermore, to improve the accuracy of search results, after obtaining the aforementioned candidate visual media, the mobile phone can perform a secondary verification of these candidate visual media to further filter them and obtain visual media that match the filtered search statement, i.e., the search statement. Optionally, the mobile phone can include visual media with a vector similarity 1 greater than threshold 1 (or the first threshold) and visual media with a vector similarity 2 greater than threshold 1 as candidate visual media. In other words, the mobile phone can use the visual media determined by the base model and the fine-tuning model as candidate visual media, i.e., it performs two searches. The process of secondary verification of candidate visual media will be described below.
[0265] S416. The multimodal understanding module sends the information of the filtered search statement and the image information of the candidate visual media to the secondary confirmation module.
[0266] The information for filtering search statements includes one or more of the following: the filtering search statement, the word segmentation results of the filtering search statement, the tag 1 included in the filtering search statement, and the text feature vector of the filtering search statement.
[0267] The image information of the candidate visual media includes one or more of the following: the similarity between the candidate visual media and the filtering search statement, the label 2 of the candidate visual media, and the image feature vector of the candidate visual media.
[0268] In some embodiments, the image feature vectors and text feature vectors described above are determined based on a multimodal model including a base model and a fine-tuning model. Correspondingly, the text feature vector of the filtered search statement may include the text feature vector of the filtered search statement corresponding to the base model and the text feature vector of the filtered search statement corresponding to the fine-tuning model. The image feature vector of the candidate visual media may include the image feature vector of the candidate visual media corresponding to the base model and the image feature vector of the candidate visual media corresponding to the fine-tuning model.
[0269] The similarity between the candidate visual media and the filtered search statement can include the similarity between the candidate visual media and the filtered search statement corresponding to the base model (or similarity 1), and the similarity between the candidate visual media and the filtered search statement corresponding to the fine-tuned model (or similarity 2).
[0270] Of course, the above image feature vectors, text feature vectors, and similarity may also be determined by only one model (such as the above base model or fine-tuned model), and the present application does not limit it.
[0271] In some embodiments, the multimodal understanding module may perform word segmentation extraction on the filtered search statement to obtain the word segmentation result of the filtered search statement. For example, if the filtered search statement is "The child is playing at the seaside", the word segmentation result is: child, is, at, seaside, and playing. Exemplarily, the multimodal understanding module may perform word segmentation extraction on the filtered search statement through the natural language understanding engine service (the natural language understanding, NLU). Alternatively, the multimodal understanding module may perform word segmentation extraction on the filtered search statement through the natural language understanding module.
[0272] After obtaining the word segmentation result of the filtered search statement, the multimodal understanding module may perform label mapping on the word segmentation result to obtain the label included in the filtered search statement, that is, label 1 (or referred to as the second label). Specifically, for each word in the word segmentation result, the filtering search module determines whether the word exists in the preset label table. When the word does not exist in the preset label table, the multimodal understanding module may determine that the word does not have a corresponding label, that is, the filtered search statement does not include the label corresponding to the word.
[0273] When the word exists in the preset label table, the multimodal understanding module may use the label corresponding to the word in the preset label table as the label 1 included in the filtered search statement. For example, the word segmentation includes "cat", and the label corresponding to "cat" in the preset label table is "cat". Therefore, the label included in the filtered search statement includes "cat". It should be noted that the words in the word segmentation result of the filtered search statement and the labels corresponding to the words may be the same or different.
[0274] In some embodiments, the above label 2 (or referred to as the first label) of the candidate visual media is obtained from the label library. The label 2 of the visual media in the label library represents the classification label of the visual media, and it may be determined by the mobile phone (such as the multimodal understanding module in the mobile phone) using an image classification model to identify the visual media. Exemplarily, the classification label of the visual media may also be determined offline by the mobile phone.
[0275] S417. The secondary confirmation module determines the Boolean value corresponding to each candidate visual media based on the information of the filtered search statement and the image information of the candidate visual media.
[0276] S418. For each candidate visual media, when the Boolean value corresponding to the candidate visual media is true, the secondary confirmation module determines that the candidate visual media is visual media 1.
[0277] In this embodiment, the multimodal understanding module inputs the information of the filtering search statement and the image information of the candidate visual media to the secondary confirmation module, so that the secondary confirmation module determines the Boolean value corresponding to each candidate visual media. The Boolean value corresponding to the candidate visual media indicates whether the candidate visual media is a visual media that matches the filtering search statement, thereby realizing the secondary confirmation of the candidate visual media.
[0278] In some embodiments, the aforementioned secondary confirmation module may include a binary classification module, a dynamic threshold module, a label confirmation module, and a whitelist threshold module. The binary classification module may employ a binary classification model to determine whether there is a correlation between the filtered search statement and the candidate visual media, thereby obtaining the Boolean value of the candidate visual media. The dynamic threshold module may employ a dynamic threshold model to determine a threshold 2 that matches the length of the filtered search statement, thereby comparing the similarity between the candidate visual media and the filtered search statement with this threshold 2 to obtain the Boolean value of the candidate visual media. The label confirmation module may compare the label 2 of the candidate visual media with the label 1 included in the filtered search statement to obtain the Boolean value of the candidate visual media. The whitelist threshold module may determine a threshold 3 that matches the entire filtered search statement based on whether the filtered search statement hits a preset dictionary, thereby comparing the similarity between the candidate visual media and the filtered search statement with this threshold 3 to obtain the Boolean value of the candidate visual media. The detailed process by which the secondary confirmation module determines the candidate visual media through the binary classification module, dynamic threshold module, label confirmation module, and whitelist threshold module, i.e., the implementation process of S417 described above, can be referred to the relevant description below, and will not be introduced here.
[0279] S419, The secondary confirmation module returns visual media 1 to the search module.
[0280] S420. The search module obtains visual media 2 based on the index of non-visual semantic subjects and visual media in the above search statement.
[0281] S421, The search module uses the intersection of visual media 1 and visual media 2 as the search result.
[0282] In this embodiment of the application, the search module searches for visual media (or visual media 2) that correspond to the filtered non-visual semantic subjects in the search statement input by the user from the constructed index. That is, it searches for visual media whose attributes match the non-visual semantic subject and uses it as visual media 2.
[0283] The search module then determines the intersection of visual media 2 and visual media 1 to obtain the target visual media. The attributes of the target visual media match the non-visual semantic subject of the search query, and the visual content of the target search media corresponds to the visual semantic subject of the search query. For example, if the search query is "sky photographed this year," then visual media 1 includes sky content, visual media 2 was captured this year, therefore the target visual media includes sky content, and the target visual media was captured this year.
[0284] S422. The search module sends the search results to the image library service module.
[0285] S423, The image gallery service module displays search results.
[0286] In this embodiment of the application, the image gallery service module can display search results to the user, for example: through the above... Figure 3A Interface 305, the above Figure 3A The search results are displayed in interface 308 as described above. Figure 3A As shown in interface 305, the search results include 239 images, but interface 305 only displays thumbnails of 8 of these images. To view all 239 images, the user can click the "More" option in interface 305. In response to this action, the phone displays... Figure 3A Interface 308.
[0287] In some embodiments, the gallery service module can sort and display target visual media based on their acquisition time in the search results. For example, the target visual media can be sorted in order of acquisition time from earliest to latest, so that those acquired earlier are displayed first.
[0288] In other embodiments, the gallery service module can sort the target visual media in the search results based on the similarity between the target visual media and the filtered search query. For example, the target visual media can be sorted in descending order of similarity, so that the target visual media with higher similarity is displayed earlier. That is, the similarity between the target visual media displayed earlier and the filtered search query is greater than or equal to the similarity between the target visual media displayed later and the filtered search query. Here, the similarity between the target visual media and the filtered search query can be similarity 1 or similarity 2, or a similarity calculated based on similarity 1 and similarity 2.
[0289] Optionally, if the filtered search statement belongs to whitelist 1 above, the similarity between the target visual media and the filtered search statement can be similarity 2, that is, the similarity between the visual media corresponding to the fine-tuning model and the filtered search statement. If the filtered search statement does not belong to whitelist 1 above, the similarity between the target visual media and the filtered search statement can be similarity 1, that is, the similarity between the visual media corresponding to the base model and the filtered search statement.
[0290] Optionally, the similarity calculated based on similarity 1 and similarity 2 can be the average of similarity 1 and similarity 2. Alternatively, a weighted sum of similarity 1 and similarity 2 can be calculated; this application does not limit the specific calculation method.
[0291] The following section will continue to describe one possible implementation process of the above S417, such as... Figure 8 As shown, the process may include S417a-S417g.
[0292] S417a, The secondary confirmation module determines whether the preset fine-tuning model vocabulary contains the filtered search statement.
[0293] In this embodiment, the secondary confirmation module can determine whether there is a filtering search statement in the preset fine-tuning model vocabulary (or the first preset vocabulary) to determine whether the filtering search statement is the branch corresponding to the base model or the branch corresponding to the fine-tuning model. That is, it determines whether to perform secondary confirmation through the module corresponding to the base model or through the module corresponding to the fine-tuning model.
[0294] If the preset fine-tuning vocabulary does not contain a filtering search statement, it indicates that the filtering search statement is the branch corresponding to the base model. In other words, it indicates that the candidate visual media is reconfirmed by the module corresponding to the base model in order to select visual media that better matches the filtering search statement from the candidate visual media. The reconfirmation module can execute S417b.
[0295] If the above-mentioned preset fine-tuning vocabulary contains a filtering search statement, it indicates that the filtering search statement is the branch corresponding to the fine-tuning model. That is, it indicates that the candidate visual media is reconfirmed by the module corresponding to the fine-tuning model in order to select the visual media that is more compatible with the filtering search statement from the candidate visual media. The reconfirmation module can execute S417f.
[0296] S417b, the secondary confirmation module inputs the text feature vector of the filtered search statement and the image feature vector of each candidate visual medium into the binary classification module, and obtains the Boolean value 1 corresponding to each candidate visual medium output by the binary classification module.
[0297] In this embodiment of the application, it is assumed that the number of candidate visual media is m. For example... Figure 9AAs shown, the secondary confirmation module inputs the image feature vectors of each of the m candidate visual media and the text feature vector of the filtering search statement into the binary classification model in the binary classification module. For each candidate visual media, the binary classification model determines the matching degree between the candidate visual media and the search statement based on the image feature vector of the candidate visual media and the text feature vector of the filtering search statement. This matching degree represents the degree of correlation between the candidate visual media and the filtering search statement; the higher the matching degree, the more relevant the candidate visual media and the filtering search statement are. Then, the binary classification model compares the matching degree between the candidate visual media and the filtering search statement with a classification threshold (or a preset classification threshold) to determine whether the candidate visual media matches the search statement, obtaining a Boolean value (bool) 1 for the candidate visual media, and outputting the Boolean value 1 for each of the m candidate visual media. If the matching degree is greater than the classification threshold, it indicates that the candidate visual media matches the filtering search statement, and the Boolean value 1 for the candidate visual media is true. If the matching degree is less than or equal to the classification threshold, it indicates that the candidate visual media does not match the filtered search statement, and the Boolean value 1 corresponding to the candidate visual media is false. Optionally, the aforementioned Boolean value 1 can also be used as the first matching value.
[0298] The binary classification model is a pre-trained model using a training sample set. It can determine whether image feature vectors and text feature vectors match. This binary classification model can be trained on the aforementioned mobile phone or other devices. The following will use the example of training the binary classification model on a mobile phone to introduce the training process of the binary classification model.
[0299] In some embodiments, if the preset fine-tuning model vocabulary contains a filtering search statement, it indicates that the filtering search statement corresponds to the branch of the fine-tuning model. Since the branch of the fine-tuning model does not include the binary classification module, the secondary confirmation module does not need to execute the above S417b.
[0300] If the word segmentation of the filtering search statement is not present in the above-mentioned preset fine-tuning model vocabulary, it indicates that the filtering search statement corresponds to the branch of the base model. The text feature vector of the filtering search statement in S417b can include the text feature vector of the filtering search statement corresponding to the base model. The image feature vector of the candidate visual media in S417b can include the image feature vector of the candidate visual media corresponding to the base model.
[0301] It is understandable that if the input parameters of the secondary confirmation module only include an image feature vector of the candidate visual media and a text feature vector of the filtering search statement, but not the image feature vectors of the candidate visual media corresponding to the fine-tuning model and the candidate visual media corresponding to the base model, then the image feature vector of the candidate visual media in S417b above is the image feature vector of the input candidate visual media, and the text feature vector of the filtering search statement is the text feature vector of the input filtering search statement.
[0302] For example, the training sample set of the binary classification model described above includes multiple training data sets, each of which includes a sample image and its corresponding descriptive text. The training sample set may include a positive sample training set (referred to as positive samples) and a negative sample training set (referred to as negative samples). The sample images in the training data included in the positive samples are matched with the descriptive text corresponding to the sample images. For example, as... Figure 9B The image shown, with the corresponding descriptive text, is of a little boy holding a basket. Because... Figure 9B The visual content depicted is a little boy holding a basket, which matches its corresponding descriptive text. Therefore, the image and its corresponding descriptive text can be used as a training data point in the positive sample.
[0303] Negative samples are training data where the sample images do not match the corresponding descriptive text. For example, ... Figure 9B The image shown has the descriptive text "flower". Because... Figure 9B The visual content expressed does not match the descriptive text. Therefore, the sample image and its corresponding descriptive text can be used as a training data point in the negative samples.
[0304] It should be noted that since users generally use natural language when searching for visual media, the descriptive text corresponding to the sample images is also in natural language to match the actual search situation.
[0305] To improve training accuracy, it's necessary to ensure a sufficient amount of training data in both positive and negative samples. Considering the low efficiency of manually obtaining positive and negative samples, mobile phones can automatically generate them. The following will use the aforementioned binary classification model, an image-text matching model, as an example to further explain the process of generating positive and negative samples. Figure 10 As shown, the process is as follows:
[0306] S1. The binary classification module obtains P sample data pairs. Each of the P sample data pairs includes a sample image and its corresponding descriptive text.
[0307] S2, the binary classification module extracts subtext from the descriptive text of each sample data.
[0308] For example, the subtext mentioned above may include descriptive phrases and / or nouns from the descriptive text. For example, subtext generally does not include adverbs, verbs, or other words in the descriptive text that do not correspond to actual visual content. The binary classification module can segment the descriptive text to identify descriptive phrases and nouns, thus obtaining the subtext within the descriptive text. For example, if the descriptive text is "a little boy sitting in a basket," the subtexts are "basket," "little boy," and "boy." Another example is "a little boy sitting in a basket sucking his thumb." After segmenting this descriptive text, the subtexts obtained are "basket," "boy," "little boy," and "little boy sucking his thumb," where "little boy sucking his thumb" can represent a descriptive phrase.
[0309] In some embodiments, the binary classification module can use the Han Language Processing (HanLP) package to segment the descriptive text corresponding to the sample images.
[0310] S3. For each sample image, the binary classification module generates a text set corresponding to the sample image based on the description text corresponding to the sample image and the subtext in the description text.
[0311] Here, the text set corresponding to the sample image represents the set of descriptions of the visual content corresponding to the sample image. The text set corresponding to the sample image can include the descriptive text corresponding to the sample image and each subtext within the descriptive text. For example, the descriptive text corresponding to the sample image is "a little boy holding a basket," and the subtexts within the descriptive text include "basket," "little boy," and "boy." Accordingly, the text set corresponding to this image is {basket, little boy, boy, little boy holding a basket}.
[0312] S4. The binary classification module randomly selects a text element from the text set corresponding to each sample image to obtain text element 1 corresponding to each sample image.
[0313] The text elements in the text collection can be subtext or descriptive text.
[0314] In some embodiments, the binary classification module can randomly select text elements from the text sets corresponding to each sample image according to a preset ratio. This preset ratio includes the proportion of selected text elements that are descriptive text and the proportion of selected text elements that are sub-text. For example, if the preset ratio is 8:2, the probability of selecting a text element that is descriptive text is 80%, and the probability of selecting a text element that is sub-text is 20%. Since the image encoder and text encoder are generally trained using the image and its corresponding text as a whole, the proportion of selected text elements that are descriptive text in the preset ratio will be greater than the proportion of selected text elements that are sub-text, ensuring the training effect of the model.
[0315] S5. For each sample image, the binary classification module takes the text element 1 corresponding to the P sample images as the sample text element corresponding to that sample image.
[0316] In this model, one sample image from a set of P sample images corresponds to P sample text elements. A sample image and its corresponding text element constitute one sample, resulting in multiple samples. Since the model needs to be trained using both positive and negative samples, it is necessary to distinguish between positive and negative samples after obtaining them, thus training a binary classification model using both positive and negative samples. The process of distinguishing between positive and negative samples will be described below.
[0317] In this embodiment, the binary classification module initially obtains P sample data pairs. All P sample data pairs are positive samples. However, training the binary classification model also requires negative samples. Therefore, the binary classification module can perform image-text matching between the P sample images and a text element (i.e., text element 1) from the corresponding text set of the P sample images. That is, for each sample image in the P sample images, the binary classification module can use all P text elements 1 as the corresponding sample text element. Each sample image and each text element 1 in the P text elements 1 form a sample, thereby obtaining P*P samples, increasing the number of samples, and thus achieving rapid sample generation. To train the binary classification model, the model needs to distinguish between positive and negative samples among the P*P samples. The process of distinguishing between positive and negative samples will be described below.
[0318] S6. For each sample text element corresponding to the sample image, the binary classification module determines whether the sample text element belongs to the text set corresponding to the sample image.
[0319] In this embodiment of the application, for each sample text element corresponding to the sample image, the binary classification module can determine whether the sample text element corresponds to the visual content of the sample image by judging whether the sample text element belongs to the text set corresponding to the sample image, thereby determining whether the sample text element and the sample image constitute a positive sample.
[0320] If the sample text element is in the text set corresponding to the sample image, it indicates that the sample text element corresponds to (i.e., matches) the visual content of the sample image. In other words, the sample consisting of the sample text element and the sample image is a positive sample. Therefore, the binary classification module can execute S7.
[0321] If the sample text element is not part of the text set corresponding to the sample image, it indicates that the sample text element does not correspond to the visual content of the image (i.e., they do not match). In other words, the sample formed by the sample text element and the sample image is not a positive sample. Therefore, the binary classification module can execute S8.
[0322] S7. The binary classification module treats the sample image and the sample text element as positive samples.
[0323] S8, the binary classification module treats the image and the text element of the sample as negative samples.
[0324] In some embodiments, after obtaining the sample text elements corresponding to each sample image, the binary classification module can generate a corresponding sample matrix. The elements in the i-th row of the sample matrix represent the sample text elements corresponding to the i-th sample image, and the elements in the j-th column of the sample matrix represent the text element 1 corresponding to the j-th sample image, which is a text element randomly selected from the text set corresponding to the j-th sample image.
[0325] For example, the sample images include image a, image b, image c, and image d. Figure 11A In the sample matrix shown, "1" represents the text element 1 corresponding to image a (i.e., the first sample image). Similarly, "6" represents the text element 1 corresponding to image d (i.e., the fourth sample image). The first row element 50 (i.e., 1, 2, 5, 6) represents the sample text element corresponding to image a. The second column element 51 (i.e., 2, 2, 2, 2) represents the text element 1 corresponding to image b.
[0326] In this sample matrix, a sample text element and the sample image corresponding to the row containing that text element constitute a sample. For example, as described above. Figure 11A The "1" in the first row and first column shown forms a sample with image a.
[0327] In this embodiment, the binary classification model generates a corresponding sample matrix by matching P sample images with P text elements 1. This allows the binary classification module to distinguish whether the sample text elements in the sample matrix belong to positive or negative samples, thus achieving rapid determination of positive and negative samples and improving the efficiency of positive and negative sample determination. Furthermore, by generating a label matrix, the comprehensiveness of image-text matching can be guaranteed, avoiding the omission of sample images or text elements 1, such as avoiding the situation where a certain sample image is not matched with a certain text element 1.
[0328] After obtaining the sample matrix, for each sample text element in the sample matrix, the binary classification module needs to determine whether the sample text element matches the sample image corresponding to the row containing the sample text element. To improve matching efficiency, the binary classification module can use the text intersection matrix 1 to match the sample matrix to obtain positive and negative samples. The text intersection matrix 1 represents the intersection of the text sets corresponding to the sample images. For example, the process of determining the text intersection matrix 1 may include:
[0329] The binary classification module determines the overlapping elements in the text sets corresponding to any two sample images. Then, the binary classification module generates the corresponding text intersection matrix 1. In this matrix, the s-th element in the t-th row represents the overlapping element in the text set between the t-th and s-th sample images.
[0330] For example, image a (such as Figure 11B The text set corresponding to image a is {1,2,3}, the text set corresponding to image b is {2,3,4}, the text set corresponding to image c is {5,6}, and the text set corresponding to image d is {6,7}. It should be understood that these numbers actually correspond to text elements (such as the subtext and descriptive text mentioned above).
[0331] Then, the binary classification module can determine the overlapping elements (i.e., the same elements) in the text sets corresponding to any two sample images from images a, b, c, and d. For example, the same elements in the text sets corresponding to images a and b are 2 and 3.
[0332] Afterwards, the binary classification module generates, as follows: Figure 11C The text intersection matrix 1 shown includes overlapping elements between the sample images and the text sets corresponding to each sample image (i.e., image a, image b, image c, and image d). Figure 11C The elements in the first row and first column of the text intersection matrix 1 shown represent the overlapping elements in the text sets between image a and image b (i.e., the text set corresponding to image a). The elements in the first row and second column represent the overlapping elements in the text sets between image a and image b. And so on. Figure 11C The element in the 4th row and 4th column represents the overlapping elements in the text set between image d and image d.
[0333] In some embodiments, after obtaining the text sets corresponding to each of the P sample images, the binary classification module can calculate the union of the text sets corresponding to each sample image. This union can include all text elements from all text sets. Then, the binary classification module can assign a number (as described above) to each text element in the union of the text sets, such that different text elements in one text set have different numbers, and the same text element in different text sets has the same number. By assigning numbers to text elements, the efficiency of image-text matching can be improved, thereby increasing the efficiency of generating positive and negative samples.
[0334] The process of determining the text intersection matrix 1 has been introduced above. The process of matching the text intersection matrix 1 with the sample matrix to determine the positive and negative samples will be introduced below.
[0335] The binary classification module intersects the text intersection matrix 1 and the sample matrix to obtain the text intersection matrix 2.
[0336] Then, for each intersection element in the text intersection matrix 2, the binary classification module determines whether the intersection element is empty.
[0337] If the intersection element is not empty, it indicates that the visual content of the intersection element matches that of the sample image corresponding to its row. Therefore, the binary classification module can confirm that the intersection element and the sample image corresponding to its row are positive samples.
[0338] If the intersection element is empty, it indicates that the visual content of the intersection element and the sample image corresponding to its row do not match. Therefore, the binary classification module can confirm that the element and the sample image corresponding to its row are negative samples, thus achieving rapid determination of positive and negative samples. The element corresponding to the sample image can also be referred to as the text corresponding to the sample image.
[0339] Optionally, the binary classification module can distinguish whether the intersection elements in the text intersection matrix 2 belong to positive or negative samples using 1 and 0, that is, to distinguish whether the sample text elements in the sample matrix belong to positive or negative samples. When the intersection elements in the text intersection matrix 2 are empty, the binary classification module can set the label corresponding to the intersection element to 0, that is, label the sample text element corresponding to the intersection element with 0.
[0340] If the intersection element in the text intersection matrix 2 is not empty, the binary classification module can set the label corresponding to that element to 1, that is, label the sample text element corresponding to that intersection element with label 1, thus obtaining the positive and negative sample label matrix, and realizing the labeling of the sample text elements in the sample matrix. The position of the intersection element in the text intersection matrix 2 is the same as the position of the sample text element corresponding to the intersection element in the sample matrix.
[0341] Based on this, the binary classification module can use the sample text element corresponding to label 0 and the image corresponding to the row where the sample text element is located as a training data in the negative sample, and use the sample text element corresponding to label 1 and the image corresponding to the row where the sample text element is located as a training data in the positive sample, so as to realize the batch determination of training data for multiple positive and negative samples.
[0342] For example, such as Figure 11D As shown, the intersection of text intersection matrix 1 and sample matrix is obtained as text intersection matrix 2. Then, the binary classification module can determine whether the intersection elements in text intersection matrix 2 are empty, thereby determining the positive and negative sample label matrix corresponding to the sample matrix. Figure 11DThe label in the first row and first column of the positive and negative sample label matrix is 1. The element in the sample matrix corresponding to this label is the sample text element "1" in the first row and first column. This "1" and image a together form a positive sample. Figure 11D The label in the second row and first column of the positive and negative sample label matrix is 0. The sample text element in the second row and first column of the sample matrix corresponding to this label is "1". This "1" and the image b together form a negative sample.
[0343] It should be noted that, generally speaking, the elements on the diagonal of the sample matrix and the sample images corresponding to their respective rows constitute the positive samples.
[0344] In this embodiment, the binary classification module obtains a text intersection matrix 2 by intersecting the text intersection matrix 1 and the sample matrix. The positive and negative samples can be quickly determined by checking whether the elements in the text intersection matrix 2 are empty, thereby improving the efficiency of determining positive and negative samples and ensuring the comprehensiveness and completeness of the positive and negative sample distinction, so as to enable the model to be trained quickly.
[0345] The process of determining positive and negative samples has been introduced above. The process of training a binary classification model using positive and negative samples will be introduced below.
[0346] S9. The binary classification module uses positive and negative samples to train the image-text matching model, resulting in the trained image-text matching model.
[0347] For example, the image-text matching model can be a multilayer perceptron (MLP) model. Figure 11E As shown, the binary classification module uses a text encoder to process the sample text elements in both positive and negative samples, obtaining the text feature vectors of the sample text elements. Simultaneously, the binary classification module uses an image encoder to process the sample images in both positive and negative samples, obtaining the image feature vectors of the sample images. Then, for each sample in both positive and negative samples, the binary classification module concatenates the text feature vectors of the sample text elements and the image feature vectors of the sample image. The binary classification module then inputs the concatenated text feature vector and image feature vector into the MLP model to train the MLP model. During training, the binary classification module can use a loss function to test the difference between the predicted values and actual values output by the trained MLP model. The predicted value indicates whether the predicted sample image matches the text. The actual value indicates the actual matching situation between the sample image and the text.
[0348] In some embodiments, the loss function described above can be a focal loss function. Specifically, the focal loss function can be FL(p t )=-at1(1-pt ) γ log(p t Here, pt represents the probability that the image-text matching model predicts a match between a sample image and its corresponding text, i.e., the probability that the predicted sample image and its corresponding text are positive samples. at1 is a factor that adjusts the weights of positive and negative samples, which can be set according to the number of positive and negative samples to adjust for imbalances in the number of positive and negative samples; for example, at1 is 0.1. γ is an adjustment factor used to reduce the loss contribution of easily distinguishable samples. It should be understood that the larger pt is, the smaller the value of the loss function.
[0349] Optionally, the image encoder described above can be a stacked autoencoder, and the text encoder can be a counting vector.
[0350] It should be noted that the focus loss function described above is only one example of a loss function. This loss function can also be other types of loss functions, such as a BCE-type loss function. Furthermore, the MLP model described above is only one example of an image-text matching model. This image-text matching model can also be other deep learning models, and this application does not limit it.
[0351] In some embodiments, the aforementioned positive and negative sample label matrix can be used when calculating the value of the loss function.
[0352] In some embodiments, the above-described image-text matching model for secondary verification of candidate visual media is merely an example; the image-text matching model can also be directly used to search for visual media. For instance, after a user enters a search query, the image-text matching model can directly utilize the text feature vector of the search query (or the filtered search query) and the image feature vectors of various visual media on the mobile phone to determine whether the visual media matches the search query and obtain the corresponding Boolean value.
[0353] The above describes the process of determining whether candidate visual media matches the filter search statement using a binary classification module. The following section will continue to describe the process of determining whether candidate visual media matches the filter search statement using a dynamic threshold module, in conjunction with S417c.
[0354] S417c, the binary classification module inputs the filtering search statement and the similarity between each candidate visual media and the filtering search statement into the dynamic threshold module, and obtains the Boolean value 2 corresponding to each candidate visual media output by the dynamic threshold module.
[0355] In this embodiment, generally, different lengths of search statements correspond to different similarity thresholds 1. The longer the search statement, the higher the similarity between the visual content of the visual media and the search statement needs to be, and correspondingly, the higher the similarity threshold 1 needs to be. Therefore, the secondary confirmation module can use the dynamic threshold module to determine the dynamic threshold that matches the length of the filtered search statement, that is, to determine the similarity threshold 1 (or the first similarity threshold) corresponding to the filtered search statement, and then use the similarity threshold 1 corresponding to the filtered search statement to determine the Boolean value 2 corresponding to each candidate visual media. Optionally, the Boolean value 2 can also be used as the second matching value.
[0356] For example, such as Figure 12 As shown, the process by which the dynamic thresholding module determines the Boolean value 2 corresponding to each candidate visual media can include: First, the dynamic thresholding module can determine the length of the filtering search statement through the filtering search statement. Then, based on the length of the filtering search statement, the dynamic thresholding model, combined with t = parameter 1 * L + parameter 2, determines the similarity threshold 1 corresponding to the filtering search statement. Here, parameter 1 and parameter 2 are pre-set parameters; for example, parameter 1 is 0.05 and parameter 2 is 0.33. Optionally, the dynamic thresholding module can also determine the length of the filtering search statement through the word segmentation results of the filtering search statement.
[0357] Subsequently, for each of the m candidate visual media, the dynamic thresholding module compares the similarity between the candidate visual media and the filter search statement with the similarity threshold 1 corresponding to the filter search statement. If the similarity between the candidate visual media and the filter search statement is less than the similarity threshold 1, it indicates that the visual content of the candidate visual media has a low degree of similarity to the filter search statement, and the dynamic thresholding module determines that the Boolean value 2 corresponding to the candidate visual media is false.
[0358] If the similarity between the candidate visual media and the filter search statement is greater than or equal to the similarity threshold 1, it indicates that the visual content of the candidate visual media is highly similar to the filter search statement, and the dynamic threshold module can determine that the Boolean value 2 corresponding to the candidate visual media is true.
[0359] Optionally, as described above Figure 12 As shown, the range of values for t can be greater than or equal to parameter 3, meaning min(t, parameter 3), and the range of values for t can be less than or equal to parameter 4, meaning max(t, parameter 4). Parameters 3 and 4 are preset values.
[0360] In some embodiments, since the dynamic threshold module corresponding to S417c belongs to the branch corresponding to the base model, the similarity between the candidate visual media and the filter search statement in S417c includes the similarity between the candidate visual media and the filter search statement corresponding to the base model (i.e., similarity 1). Accordingly, the dynamic threshold module can determine whether the similarity 1 between the candidate visual media and the filter search statement is less than the similarity threshold 1. If the similarity 1 is greater than or equal to the similarity threshold 1, the dynamic threshold module can determine that the Boolean value 2 corresponding to the candidate visual media is true. If the similarity 1 is less than the similarity threshold 1, the dynamic threshold module can determine that the Boolean value 2 corresponding to the candidate visual media is false.
[0361] Understandably, if the input parameters of the secondary confirmation module only include a similarity between the candidate visual media and the filter search statement, and do not include the similarity between the candidate visual media and the filter search statement corresponding to the fine-tuning model and the candidate visual media and the filter search statement corresponding to the base model, then the image feature vector of the candidate visual media in S417c above is the similarity between the input candidate visual media and the filter search statement.
[0362] In addition, the similarity between the aforementioned candidate visual media and the filtered search statement can also be determined by other multimodal models, and this application does not limit it.
[0363] The above describes the process of using the dynamic threshold module to determine the Boolean value corresponding to the candidate visual media. The following section will continue to introduce the process of using the label confirmation module to determine whether the candidate visual media matches the filter search statement, in conjunction with S417d.
[0364] S417d, the secondary confirmation module inputs the word segmentation result of the above filtering search statement, the tag 1 included in the filtering search statement, and the tag 2 of each candidate visual medium into the tag confirmation module, and obtains the Boolean value 3 corresponding to each candidate visual medium output by the tag confirmation module.
[0365] In this embodiment, the secondary confirmation module can use the tag confirmation module to determine whether the tag 2 of the candidate visual media exists in the word segmentation of the filtering search statement or the tag 1 included in the filtering search statement, so as to determine whether the candidate visual media matches the filtering search statement, thereby determining the Boolean value 3 corresponding to the candidate visual media. Optionally, the Boolean value 3 can also be used as a third matching value.
[0366] For example, such as Figure 13 As shown, the process by which the label confirmation model determines the Boolean value 3 corresponding to the candidate visual media can include:
[0367] First, for each of the m candidate visual media, the tag confirmation module can determine whether the tag 2 of the candidate visual media includes the word segmentation result of the filtered search statement, and determine whether the tag 2 of the candidate visual media includes the tag 1 included in the filtered search statement, that is, determine whether the candidate visual media includes the visual content corresponding to the filtered search statement.
[0368] If the candidate visual media's tag 2 includes at least one word of the filtered search query, or if the candidate visual media's tag 2 includes at least one tag 1 included in the filtered search query, it indicates that the candidate visual media matches the filtered search query. The tag confirmation module can determine that the candidate visual media may be the visual media needed by the user. Therefore, the tag confirmation module can determine the Boolean value 3 corresponding to the candidate visual media as true. For example, if the filter search query's tag 1 includes "child" and "flower," and the filter search query's word segmentation results include the words "child," "holding," and "flower," then if the candidate visual media's tag 2 includes "child," "holding," or "flower," or if the candidate visual media's tag 2 includes "child" or "flower," then the Boolean value 3 corresponding to the candidate visual media is determined to be true.
[0369] If the tag 2 of the above candidate visual media does not include all the words of the filter search statement, and the tag 2 of the candidate visual media does not include all the tags 1 included in the filter search statement, it indicates that the candidate visual media does not match the filter search statement, and the candidate visual media may not be the visual media required by the user. Therefore, the tag confirmation module can determine the Boolean value 3 corresponding to the candidate visual media as false, so as to obtain the Boolean value 3 of m candidate visual media.
[0370] Optionally, to improve the accuracy of search results, the tag confirmation module may determine that the Boolean value 3 corresponding to the candidate visual media is true if the tag 2 of the candidate visual media includes all the words of the filtered search statement, or includes all the tags 1 included in the filtered search statement.
[0371] If the candidate visual media's tag 2 does not include at least one word of the filtering search statement, and the filtering search statement does not include at least one tag 1, then the Boolean value 3 corresponding to the candidate visual media is determined to be false.
[0372] The above describes the process by which the secondary confirmation module uses the tag confirmation module to determine the Boolean value 3 corresponding to each candidate visual media. The following will continue to describe the process by which the secondary confirmation module uses the whitelist threshold module to determine whether the candidate visual media matches the filtering search statement.
[0373] S417e, the secondary confirmation module inputs the similarity between the filtered search statement, the candidate visual media and the filtered search statement into the whitelist threshold module, and obtains the Boolean value 4 corresponding to each candidate visual media output by the whitelist threshold module.
[0374] In this embodiment, the secondary confirmation module can determine whether the filtered search statement hits the dictionary through the whitelist threshold module, thereby determining whether a similarity threshold 2 corresponding to the filtered search statement exists, and accurately determining the similarity threshold. For example, the whitelist threshold module can determine whether the dictionary contains the filtered search statement. If it does, the whitelist threshold module takes effect, using the similarity threshold corresponding to the filtered search statement in the dictionary as the similarity threshold 2, and thus using the similarity threshold 2 to filter candidate visual media matching the filtered search statement. If it does not exist, the whitelist threshold module does not take effect, meaning there is no need to use the whitelist threshold module to determine the Boolean value 4 corresponding to each candidate visual media. The dictionary includes a key and its corresponding value. The key represents a preset search text, and the value is the similarity threshold corresponding to that preset search text. The key in the dictionary can be search text frequently entered by the user, i.e., a pre-tested search statement. Optionally, the Boolean value 4 can also be used as a fourth matching value.
[0375] In some embodiments, the branches corresponding to the fine-tuning model and the base model have corresponding dictionaries. For example... Figure 14A As shown, when the filtering search statement is the branch corresponding to the base model, the whitelist threshold module can determine whether the dictionary (or second preset vocabulary) corresponding to the base model includes the filtering search statement, that is, whether the dictionary corresponding to the base model has the same key as the filtering search statement.
[0376] If the dictionary corresponding to the base model includes the filter search statement, it indicates that the filter search statement hits the dictionary corresponding to the base model, that is, the dictionary corresponding to the base model contains the same key as the filter search statement. Then the whitelist threshold module can use the value corresponding to the filter search statement in the dictionary corresponding to the base model as the similarity threshold 2.
[0377] Subsequently, for each of the m candidate visual media, the whitelist threshold module compares the similarity between the candidate visual media corresponding to the base model and the filtering search threshold with a similarity threshold 2. If the similarity between the candidate visual media corresponding to the base model and the filtering search threshold is less than the similarity threshold 2, the whitelist threshold module determines that the corresponding Boolean value 4 for the candidate visual media is false. If the similarity is greater than or equal to the similarity threshold 2, it indicates that the visual content of the candidate visual media is highly similar to the filtering search statement, and the whitelist threshold module determines that the corresponding Boolean value for the candidate visual media is true.
[0378] Accordingly, when the filtering search statement corresponds to the branch of the base model, S418 above can be: for each candidate visual medium, if the Boolean value corresponding to the candidate visual medium is true, the secondary confirmation module (such as the fusion module in the secondary confirmation module) can treat the candidate visual medium as visual medium 1. The Boolean values corresponding to the candidate visual medium include Boolean value 1, Boolean value 2, Boolean value 3, and Boolean value 4. That is, the secondary confirmation module can take the union of Boolean values 1, 2, 3, and 4 corresponding to the candidate visual medium to obtain the Boolean value corresponding to the candidate visual medium.
[0379] If all Boolean values corresponding to the candidate visual medium are false, it means that Boolean values 1, 2, 3 and 4 corresponding to the candidate visual medium are all false, and the secondary confirmation module (such as the fusion module in the secondary confirmation module) can choose not to use the candidate visual medium as visual medium 1.
[0380] It should be noted that since both the whitelist threshold module and the dynamic threshold module confirm the similarity threshold corresponding to the filtering search statement, the similarity threshold determined by the whitelist threshold module is more accurate. Therefore, when the whitelist threshold module is in effect, that is, when the Boolean value corresponding to the candidate visual media is determined using the similarity threshold 2 corresponding to the filtering search statement, the dynamic threshold module may not be in effect. The secondary confirmation module may not use the similarity threshold 1 corresponding to the filtering search statement to determine the Boolean value corresponding to the candidate visual media, thus avoiding unnecessary determination of similarity thresholds and unnecessary screening of candidate visual media.
[0381] In some embodiments, the binary classification module, dynamic threshold module, label confirmation module, and whitelist threshold module in the secondary confirmation module can determine the Boolean value corresponding to the candidate visual media in parallel or sequentially. Regardless of whether the determination is parallel or sequential, if one module in the secondary confirmation module determines that the Boolean value corresponding to the candidate visual media is true, other modules do not need to continue determining the Boolean value corresponding to that candidate visual media. That is, there is no need to input the relevant information of that candidate visual media into other models, avoiding unnecessary Boolean value determination and ensuring low transmission costs.
[0382] Furthermore, the aforementioned secondary confirmation module, including a binary classification module, a dynamic threshold module, a label confirmation module, and a whitelist threshold module, is merely an example. The secondary confirmation module can include one or more of these modules to improve the search efficiency for visual media. For instance, if the secondary confirmation module includes one of these modules, such as a binary classification module, a dynamic threshold module, a label confirmation module, and a whitelist threshold module, and this module includes a binary classification module, the corresponding Boolean value for the candidate visual media can include Boolean value 1. As another example, if the secondary confirmation module includes two of these modules, such as a binary classification module, a dynamic threshold module, a label confirmation module, and a whitelist threshold module, and this module includes both a binary classification module and a dynamic threshold module, the corresponding Boolean value for the candidate visual media can include Boolean value 1 and Boolean value 2. As yet another example, if the secondary confirmation module includes three of these modules, such as a binary classification module, a dynamic threshold module, a label confirmation module, and a whitelist threshold module, and this module includes a binary classification module, a dynamic threshold module, and a label confirmation module, the corresponding Boolean value for the candidate visual media can include Boolean value 1, Boolean value 2, and Boolean value 3.
[0383] Furthermore, the content included in the aforementioned candidate visual media image information and filtering search statement information, i.e., the input parameters of the secondary confirmation module, is merely an example and can be adaptively set according to the modules included in the secondary confirmation module. For example, the secondary confirmation module includes a binary classification model, the aforementioned candidate visual media image information may include the image feature vectors of the candidate visual media, and the aforementioned filtering search statement information may include the text feature vectors of the filtering search statement.
[0384] The above describes how, when the filtered search statement corresponds to a branch of the base model, the secondary confirmation module can sequentially use the binary classification module, dynamic threshold module, label confirmation module, and whitelist threshold module to determine the Boolean value corresponding to the candidate visual media. The following will continue to describe how, when the filtered search statement corresponds to a branch of the fine-tuning model, the secondary confirmation module can sequentially use the label confirmation module and whitelist threshold module to determine the Boolean value corresponding to the candidate visual media.
[0385] S417f, the secondary confirmation module inputs the word segmentation result of the above filtering search statement, the tag 1 included in the filtering search statement, and the tag 2 of each candidate visual media to the tag confirmation module, and obtains the Boolean value 5 corresponding to each candidate visual media output by the tag confirmation module.
[0386] The implementation process of S417f can be referred to the implementation process of S417d described above, and will not be repeated here. Optionally, the Boolean value 5 can also be used as the fifth matching value.
[0387] S417g, the secondary confirmation module inputs the similarity between the filtered search statement, the candidate visual media and the filtered search statement into the whitelist threshold module, and obtains the Boolean value 6 corresponding to each candidate visual media output by the whitelist threshold module.
[0388] The implementation process of S417g can be referenced from the implementation process of S417e described above. For example... Figure 14B As shown, when the filtering search statement corresponds to the branch of the fine-tuning model, the whitelist threshold module can determine whether the dictionary corresponding to the fine-tuning model includes the filtering search statement, that is, whether the dictionary (or third preset vocabulary) corresponding to the fine-tuning model has the same key as the filtering search statement, thereby determining whether the whitelist threshold module is effective.
[0389] After the whitelist threshold module takes effect, the secondary confirmation module can compare the similarity between the candidate visual media corresponding to the fine-tuning model and the filtered search statement with the value corresponding to the filtered search statement in the dictionary corresponding to the fine-tuning model to determine the Boolean value corresponding to the candidate visual media. Optionally, the Boolean value 6 can also be used as the sixth matching value.
[0390] Accordingly, when the filtering search statement corresponds to the branch of the fine-tuning model, S418 above can be: for each candidate visual medium, if the boolean value corresponding to the candidate visual medium is true, the secondary confirmation module (such as the fusion module in the secondary confirmation module) can treat the candidate visual medium as visual medium 1. The boolean values corresponding to the candidate visual medium include the aforementioned boolean value 5 and boolean value 6. That is, the secondary confirmation module can take the union of the boolean values 5 and 6 corresponding to the candidate visual medium to obtain the boolean value corresponding to the candidate visual medium.
[0391] In this embodiment, after obtaining candidate visual media, the mobile phone uses a secondary confirmation module to further determine the Boolean value corresponding to the candidate visual media in order to determine whether the candidate visual media matches the filter search statement. In this way, candidate visual media that matches the filter search statement can be selected from the candidate visual media to obtain the corresponding search results, ensuring the accuracy of the search results and thus ensuring user satisfaction.
[0392] It should be noted that the operations performed by the above-described module or model are merely examples, and the operations performed by the above-described module could also be performed by other modules in the mobile phone; this application does not limit them. Furthermore, the operations performed by the above-described module or model are actually performed by the mobile phone itself.
[0393] In some embodiments, when determining candidate visual media or performing secondary confirmation on candidate visual media, the mobile phone may not first filter non-visual semantic subjects in the search statement, but directly use the text feature vector of the search statement to determine candidate visual media or perform secondary confirmation on candidate media.
[0394] In some embodiments, the search and storage of the aforementioned visual media are carried out under the authorization of the user, including but not limited to notifying and reminding the user to read the relevant user agreement (notification) and sign the agreement (authorization) which includes the authorization of relevant user information before the user uses the function.
[0395] In some embodiments, this application provides a computer storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform the method described above.
[0396] In some embodiments, this application provides a computer program product that, when run on an electronic device, causes the electronic device to perform the method described above.
[0397] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated.
[0398] It should be understood that the term "embodiment" used throughout the specification means that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, various embodiments throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0399] It should also be understood that in this application, “when…”, “if” and “if” all refer to the UE or base station taking corresponding actions under certain objective circumstances, and are not time-limited, nor do they require the UE or base station to perform a judgment action, nor do they imply any other limitations.
[0400] Those skilled in the art will understand that the various numerical designations such as "first," "second," etc., involved in this application are merely for the convenience of description and are not intended to limit the scope of the embodiments of this application, nor do they indicate the order of sequence.
[0401] In this application, the use of singular pronouns to denote "one or more" rather than "one and only one," unless otherwise specified. In this application, unless otherwise specified, "at least one" is intended to mean "one or more," and "more than" is intended to mean "two or more."
[0402] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. Here, A can be singular or plural, and B can be singular or plural.
[0403] In this document, the terms "at least one of..." or "at least one of..." refer to all or any combination of the listed items. For example, "at least one of A, B, and C" can mean: A exists alone, B exists alone, C exists alone, A and B exist simultaneously, B and C exist simultaneously, and A, B, and C exist simultaneously. A can be singular or plural, B can be singular or plural, and C can be singular or plural.
[0404] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0405] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0406] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0407] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0408] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0409] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0410] The same or similar parts between the various embodiments in this application can be referred to mutually. In the various embodiments of this application, and in the various implementation methods / methods / implementations within each embodiment, unless otherwise specified or logically conflicting, the terminology and / or descriptions between different embodiments and between the various implementation methods / methods / implementations within each embodiment are consistent and can be mutually referenced. The technical features in different embodiments and the various implementation methods / methods / implementations within each embodiment can be combined according to their inherent logical relationships to form new embodiments, implementation methods, methods, or implementation approaches. The above-described embodiments of this application do not constitute a limitation on the scope of protection of this application.
[0411] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of protection of the claims. In conclusion, the above description is merely a preferred embodiment of the technical solution of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A visual media search method, applicable to electronic devices, characterized in that, The method includes: The first interface is displayed; the first interface includes a search box; The system receives a search statement entered in the search box, performs non-visual semantic subject filtering on the search statement to obtain a filtered search statement; wherein, the non-visual semantic subject does not correspond to visual content. Candidate visual media are identified; wherein, the candidate visual media are visual media whose similarity to the filtered search statement is greater than a first threshold, and the similarity is obtained based on the text feature vector of the filtered search statement and the image feature vector of the visual media in the electronic device; if the filtered search statement matches the whitelist, the text feature vector and the image feature vector are generated by a fine-tuning model; if the filtered search statement does not match the whitelist, the text feature vector and the image feature vector are generated by a base model; the fine-tuning model is obtained by training the base model on the training set of the preset search scenario corresponding to the whitelist, and the fine-tuning model and the base model reuse the text encoder and the image encoder; Based on the image information of the candidate visual media and the information of the filtering search statement, a matching value corresponding to the candidate visual media is determined; wherein, if the first preset vocabulary does not include the filtering search statement, the matching value corresponding to the candidate visual media is determined based on the first target matching value corresponding to the candidate visual media; if the first preset vocabulary includes the filtering search statement, the matching value corresponding to the candidate visual media is determined based on the second target matching value corresponding to the candidate visual media. The target visual media is determined based on the matching value corresponding to the candidate visual media; The search results are displayed, and the search results correspond to the target visual media.
2. The method according to claim 1, characterized in that, The first target matching value includes at least one of a first matching value, a second matching value, a third matching value, and a fourth matching value; the first matching value is determined based on the degree of matching between the candidate visual media and the filtered search statement; The second matching value is determined by comparing the similarity between the candidate visual media and the filtered search statement with a first similarity threshold corresponding to the filtered search statement; The third matching value is determined based on whether the visual content in the candidate visual media includes the visual content corresponding to the filtering search statement; The fourth matching value is determined by comparing the similarity between the candidate visual media and the filtered search statement with the second similarity threshold corresponding to the filtered search statement; The second target matching value includes the fifth matching value and / or the sixth matching value; The fifth matching value is determined based on whether the visual content in the candidate visual media includes the visual content corresponding to the search statement, and the sixth matching value is determined by comparing the similarity between the candidate visual media and the filtered search statement with the third similarity threshold corresponding to the filtered search statement.
3. The method according to claim 2, characterized in that, The image information of the candidate visual media includes the image feature vector of the candidate visual media, and the information of the filtering search statement includes the text feature vector of the filtering search statement; The step of determining the matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the filtering search statement includes: The image feature vectors of each candidate visual medium and the text feature vector of the filtering search statement are input into a binary classification model to obtain a first matching value corresponding to each candidate visual medium. The binary classification model is used to determine the matching degree between the candidate visual medium and the filtering search statement based on the image feature vectors of the candidate visual medium and the text feature vectors of the search statement, and to determine the first matching value corresponding to the candidate visual medium according to the matching degree and a preset classification threshold.
4. The method according to claim 3, characterized in that, The image feature vector of the candidate visual media includes the image feature vector of the candidate visual media corresponding to the base model, and the text feature vector of the filtering search statement includes the text feature vector of the filtering search statement corresponding to the base model.
5. The method according to any one of claims 2 to 4, characterized in that, The image information of the candidate visual media includes the similarity between the candidate visual media and the filtering search statement; The step of determining the matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the filtering search statement includes: Based on the length of the filtering search statement, a first similarity threshold corresponding to the filtering search statement is determined; If the similarity between the candidate visual media and the filtered search statement is greater than or equal to the first similarity threshold, the second matching value corresponding to the candidate visual media is determined to be true. If the similarity between the candidate visual media and the filtered search statement is less than the first similarity threshold, the second matching value corresponding to the candidate visual media is determined to be false.
6. The method according to claim 5, characterized in that, The similarity between the candidate visual media and the filtered search statement includes the similarity between the candidate visual media and the filtered search statement corresponding to the base model.
7. The method according to any one of claims 2 to 4, characterized in that, The image information of the candidate visual media includes a first tag of the candidate visual media; the first tag indicates the category to which the visual content included in the candidate visual media belongs; the information of the filtering search statement includes word segmentation and a second tag of the filtering search statement; the second tag of the filtering search statement is obtained by mapping the word segmentation of the filtering search statement. The step of determining the matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the filtering search statement includes: If the first tag of the candidate visual media includes a word segment of the filtering search statement, or includes a second tag of the filtering search statement, then the third matching value corresponding to the candidate visual media is determined to be true. If the first tag of the candidate visual media does not include the word segment of the filtering search statement, and the second tag of the filtering search statement does not include the first tag of the candidate visual media, then the third matching value corresponding to the candidate visual media is determined to be false.
8. The method according to any one of claims 2 to 4, characterized in that, The image information of the candidate visual media includes the similarity between the candidate visual media and the filtering search statement, and the information of the filtering search statement includes the filtering search statement itself. The step of determining the matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the filtering search statement includes: If the second preset vocabulary includes the filtering search statement, and the similarity between the candidate visual media and the filtering search statement is greater than or equal to the second similarity threshold, then the fourth matching value corresponding to the candidate visual media is determined to be true; the second similarity threshold refers to the similarity threshold corresponding to the filtering search statement in the second preset vocabulary. If the similarity between the candidate visual media and the filtered search statement is less than the second similarity threshold, the fourth matching value corresponding to the candidate visual media is determined to be false.
9. The method according to claim 2, characterized in that, The image information of the candidate visual media includes the first tag of the candidate visual media; the information of the filtering search statement includes the word segmentation and second tag of the filtering search statement; The step of determining the matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the filtering search statement includes: If the first tag of the candidate visual media includes a word segment of the filtering search statement, or includes a second tag of the filtering search statement, then the fifth matching value corresponding to the candidate visual media is determined to be true. If the first tag of the candidate visual media does not include the word segment of the filtering search statement, and the second tag of the filtering search statement does not include the first tag of the candidate visual media, then the fifth matching value corresponding to the candidate visual media is determined to be false.
10. The method according to claim 2 or 9, characterized in that, The image information of the candidate visual media includes the similarity between the candidate visual media and the filtering search statement, and the information of the filtering search statement includes the filtering search statement itself. The step of determining the matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the filtering search statement includes: If the third preset vocabulary includes the filtered search statement, and the similarity between the candidate visual media and the filtered search statement is greater than or equal to the third similarity threshold, then the sixth matching value corresponding to the candidate visual media is determined to be true. The third similarity threshold refers to the similarity threshold corresponding to the filtered search statement in the third preset vocabulary; If the similarity between the candidate visual media and the filtered search statement is less than the third similarity threshold, the sixth matching value corresponding to the candidate visual media is determined to be false.
11. The method according to claim 10, characterized in that, The similarity between the candidate visual media and the filtered search statement refers to the similarity between the candidate visual media and the filtered search statement corresponding to the fine-tuning model; the text feature vector of the filtered search statement refers to the text feature vector of the filtered search statement corresponding to the fine-tuning model.
12. The method according to any one of claims 1 to 4, characterized in that, The step of determining the target visual media based on the matching value corresponding to the candidate visual media includes: If the matching value corresponding to the candidate visual medium is true, then the candidate visual medium is determined to be the target visual medium.
13. The method according to any one of claims 2 to 4, characterized in that, If the first target matching value corresponding to the candidate visual media includes true, then the matching value corresponding to the candidate visual media is true; If the first target matching value corresponding to the candidate visual media does not include true, the matching value corresponding to the candidate visual media is false.
14. The method according to any one of claims 1 to 4, characterized in that, The image information of the candidate visual media includes at least one of the following: the similarity between the candidate visual media and the filtering search statement, the image feature vector of the candidate visual media, and the first label of the candidate visual media. The information of the filtering search statement includes at least one of the filtering search statement, the text feature vector of the filtering search statement, the second tag included in the filtering search statement, and the word segmentation of the filtering search statement.
15. The method according to any one of claims 1 to 4, characterized in that, The determination of candidate visual media includes: For each visual medium on the electronic device, the similarity between the visual medium and the filter search statement is determined based on the image feature vector of the visual medium and the text feature vector of the filter search statement. If the similarity is greater than the first threshold, the visual media is selected as the candidate visual media.
16. The method according to any one of claims 1 to 4, characterized in that, The matching value corresponding to the candidate visual media is determined based on the image information of the candidate visual media and the information of the filtering search statement.
17. An electronic device, characterized in that, The electronic device includes a display screen, a memory, and one or more processors; the display screen, the memory, and the processors are coupled; the display screen is used to display an image generated by the processor, the memory is used to store computer program code, the computer program code including computer instructions; when the processor executes the computer instructions, the electronic device performs the method as described in any one of claims 1 to 16.
18. A computer storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 16.