Visual media searching method and electronic equipment
By conducting preliminary search and secondary confirmation of visual media in electronic devices, using similarity and feature vector matching values, the problem of inaccurate visual media search under complex search statements is solved, and high-accurate visual media search is achieved.
Patent Information
- Application Number
- CN202311584703.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-11-23
AI Technical Summary
In the prior art, when processing complex search statements, it is difficult to accurately retrieve visual media, resulting in inaccurate search results.
By realizing preliminary search and secondary confirmation of visual media in electronic devices, the target visual media is determined and the accuracy of search results is improved by using the similarity between visual media and search statements and the matching value between image feature vectors and text feature vectors.
It realizes accurate retrieval of complex search statements, improves the accuracy of visual media search results, and meets users' search needs.
Smart Images

Figure CN120067394A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a visual media search method and an electronic device. Background Art
[0002] With the development of electronic devices (such as mobile phones), the shooting function of mobile phones has also developed rapidly. More and more users use mobile phones to take pictures, videos, etc., and store the taken pictures and videos in the picture gallery of the mobile phone. In addition, users can also store the pictures and screenshots downloaded by the mobile phone in the picture gallery of the mobile phone.
[0003] When a user wants to search for visual media (such as pictures, videos, etc.), the user can enter a search statement on the mobile phone. For example, for pictures taken on September 1st, the mobile phone responds to the search statement entered by the user, retrieves the pictures taken on September 1st, and obtains corresponding search results. However, the mobile phone's ability to understand search statements is limited. When the search statement is relatively complex, the mobile phone may not be able to accurately obtain the corresponding search results. Summary of the Invention
[0004] In view of this, this application provides a visual media search method and an electronic device for improving the accuracy of search results.
[0005] In a first aspect, this application provides a visual media search method applied to an electronic device. The electronic device can display a first interface, and the first interface includes a search box. The electronic device can receive a search statement input by the user in the search box. After that, the electronic device can determine candidate visual media, where the candidate visual media represents visual media on the electronic device with a similarity greater than a first threshold to the search statement. After that, for each candidate visual media, the electronic device can determine a matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the search statement.
[0006] After that, the electronic device can determine a target visual media from the candidate visual media according to the matching values corresponding to each candidate visual media. After that, the electronic device displays search results, and the search results correspond to the target visual media.
[0007] In this application, after the electronic device receives a search statement about visual media input by the user, it indicates that it needs to search for visual media that matches the search statement. The electronic device can use visual media with a similarity greater than a first threshold between the visual media and the search statement as candidate visual media to achieve a preliminary retrieval of visual media. After that, the electronic device can perform a secondary confirmation on the candidate visual media, that is, the electronic device can determine a matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the search statement, so that the electronic device can further determine whether the candidate visual media matches the search statement according to the matching value corresponding to the candidate visual media, thereby determining whether the candidate visual media can be used as the target visual media in the search results, improving the accuracy of determining the target visual media, and further improving the accuracy of determining the search results. And by the similarity between the visual media and the search statement, and determining whether the visual media is the target visual media based on the image information of the visual media and the information of the search statement, accurate search of the target visual media can be achieved regardless of whether the search statement is complex, meeting the user's search needs.
[0008] In a possible design, the image information of the above candidate visual media includes at least one of the similarity between the candidate visual media and the search statement, the image feature vector of the candidate visual media, and the first label of the candidate visual media;
[0009] The information of the search statement includes at least one of the search statement, the text feature vector of the search statement, the second label included in the search statement, and the word segmentation of the search statement.
[0010] Among them, the first label of the above candidate visual media represents the classification to which the visual content included in the candidate visual media belongs.
[0011] The second label of the above search statement is obtained by mapping the word segmentation of the search statement. Specifically, for each word segmentation of the visual media, the electronic device can determine whether the word segmentation exists in the preset label table. When the word segmentation exists in the preset label table, the label corresponding to the word segmentation in the preset label table is used as the second label.
[0012] Among them, optionally, the similarity between the above candidate visual media and the search statement may include the similarity between the candidate visual media corresponding to the base model and the search statement (that is, the similarity between the candidate visual media determined by the base model and the search statement) and / or the similarity between the candidate visual media corresponding to the fine-tuning model and the search statement (that is, the similarity between the candidate visual media determined by the fine-tuning model and the search statement).
[0013] The image feature vector of the above candidate visual media may include the image feature vector of the candidate visual media corresponding to the base model and / or the image feature vector of the candidate visual media corresponding to the fine-tuning model.
[0014] The text feature vector of the above search statement may include the text feature vector of the search statement corresponding to the base model and / or the text feature vector of the search statement corresponding to the fine-tuning model.
[0015] The above fine-tuning model is trained on the base model based on the training set corresponding to the preset search scenario, and the fine-tuning model has a corresponding support range. When the search statement falls within this support range, the fine-tuning model can accurately retrieve visual media that matches the search statement.
[0016] In a possible design, the process of determining the candidate visual media may include:
[0017] For each visual media on the electronic device, the electronic device can determine the similarity between the visual media and the search statement based on the image feature vector of the visual media and the text feature vector of the search statement.
[0018] When the similarity between the visual media and the search statement is greater than or equal to the first threshold, the electronic device determines that the visual media is a candidate visual media.
[0019] When the similarity between the visual media and the search statement is less than the first threshold, the electronic device determines that the visual media is not a candidate visual media.
[0020] In this application, the electronic device can calculate the vector similarity between the image feature vector of the visual media and the text feature vector of the search statement, that is, determine the similarity between the visual media and the search statement based on the image feature vector of the visual media and the text feature vector of the search statement. After that, the electronic device can use the similarity to determine whether the visual media is highly relevant to the search statement, so as to realize the determination of the candidate visual media and ensure the accuracy of the initial retrieval.
[0021] Optionally, when the similarity between the visual media corresponding to the base model and the search statement is greater than or equal to the first threshold, the electronic device determines that the visual media is a candidate visual media.
[0022] And / or,
[0023] When the similarity between the visual media corresponding to the fine-tuning model and the search statement is greater than or equal to the first threshold, the electronic device determines that the visual media is a candidate visual media.
[0024] In this application, the electronic device can determine the similarity between the visual media and the search statement through the base model and / or the fine-tuning model, so as to use the similarity to determine the candidate visual media. Since the base model has strong generalization ability, it can accurately judge the similarity between the visual content in the visual media and the search statement in most search scenarios. The fine-tuning model can accurately identify the similarity between the visual content in the visual media and the search statement in a specific search scenario.
[0025] In a possible design, the above-mentioned candidate visual media includes visual media corresponding to the base model with a similarity greater than or equal to the first threshold between the visual media and the search statement, and visual media corresponding to the fine-tuning model with a similarity greater than or equal to the first threshold between the visual media and the search statement, so as to ensure the comprehensiveness of the search results.
[0026] Optionally, the threshold for comparing the similarity between the visual media corresponding to the fine-tuning model and the search statement to determine the candidate visual media may not be the first threshold, but the second threshold. That is, when the similarity between the visual media corresponding to the fine-tuning model and the search statement is greater than or equal to the second threshold, the electronic device can determine that the visual media is a candidate visual media.
[0027] In a possible design, the similarity between the visual media corresponding to the base model and the search statement is determined based on the image feature vector of the visual media corresponding to the base model and the text feature vector of the visual media corresponding to the base model.
[0028] The similarity between the visual media corresponding to the fine-tuning model and the search statement is determined based on the image feature vector of the visual media corresponding to the fine-tuning model and the text feature vector of the visual media corresponding to the fine-tuning model.
[0029] In a possible design, the process of determining the matching value corresponding to the above-mentioned candidate visual media may include:
[0030] When the first preset vocabulary does not include the search statement, indicating that the search statement does not hit the whitelist corresponding to the fine-tuning model, the branch corresponding to the base model of the search statement, the electronic device can determine the matching value corresponding to the candidate visual media according to the first target matching value corresponding to the candidate visual media. Among them, the first target matching value includes at least one of the first matching value, the second matching value, the third matching value, and the fourth matching value; the first matching value is determined according to the matching degree between the candidate visual media and the search statement, that is, the first matching value is determined by the binary classification module.
[0031] The second matching value is determined by comparing the similarity between the candidate visual media and the search statement and the first similarity threshold corresponding to the search statement, that is, the second matching value is determined by the dynamic threshold module.
[0032] The third matching value is determined based on whether the visual content in the candidate visual media includes the visual content corresponding to the search statement, that is, the third matching value is determined by the label confirmation module.
[0033] The fourth matching value is determined by comparing the similarity between the candidate visual media and the search statement with the second similarity threshold corresponding to the search statement, that is, the fourth matching value is determined by the whitelist threshold module.
[0034] When the first preset vocabulary includes the search statement, it indicates that the search statement hits the whitelist corresponding to the fine-tuning model, and the electronic device can determine the matching value corresponding to the candidate visual media according to the second target matching value corresponding to the candidate visual media. Among them, the second target matching value includes the fifth matching value and / or the sixth matching value.
[0035] The fifth matching value is determined based on whether the visual content in the candidate visual media includes the visual content corresponding to the search statement, that is, the fifth matching value is determined by the label confirmation module
[0036] The sixth matching value is determined by comparing the similarity between the candidate visual media and the search statement with the third similarity threshold corresponding to the search statement, that is, the sixth matching value is determined by the whitelist threshold module.
[0037] In this application, by determining whether the search statement hits the first preset vocabulary, that is, whether it hits the whitelist corresponding to the fine-tuning model, the module used for secondary confirmation is determined, ensuring the accuracy of secondary confirmation.
[0038] In a possible design, the above matching value is a boolean value. The process of determining the target visual media according to the matching value corresponding to the candidate visual media may include:
[0039] When the matching value corresponding to the above candidate visual media is true, the electronic device uses the candidate visual media as the target visual media.
[0040] When the matching value corresponding to the above candidate visual media is false, the electronic device determines that the candidate visual media is not the target visual media.
[0041] In a possible design, when the first preset vocabulary does not include the search statement, the matching value corresponding to the candidate visual media is the union of the matching values included in the first target matching value corresponding to the candidate visual media. Specifically, when the first target matching value corresponding to the candidate visual media includes true, the electronic device can determine that the matching value corresponding to the candidate visual media is true. When the first target matching value corresponding to the candidate visual media does not include true, that is, all the matching values in the first target matching value are false, the electronic device can determine that the matching value corresponding to the candidate visual media is false.
[0042] Optionally, the first target matching value may include a first matching value, a second matching value, a third matching value, and a fourth matching value. Correspondingly, when the first matching value, the second matching value, the third matching value, and the fourth matching value are all false, the electronic device can determine that the matching value corresponding to the candidate visual media is false.
[0043] When the first matching value, the second matching value, the third matching value, or the fourth matching value is true, the electronic device can determine that the matching value corresponding to the candidate visual media is true.
[0044] In a possible design, when the first preset vocabulary includes the search statement, the matching value corresponding to the candidate visual media is the union of the matching values included in the second target matching value corresponding to the candidate visual media. Specifically, when the second target matching value corresponding to the candidate visual media includes true, the electronic device can determine that the matching value corresponding to the candidate visual media is true.
[0045] When the second target matching value corresponding to the candidate visual media does not include true, that is, all the matching values in the second target matching value are false, the electronic device can determine that the matching value corresponding to the candidate visual media is false.
[0046] Optionally, the first target matching value may include a fifth matching value and a sixth matching value. Correspondingly, when the fifth matching value and the sixth matching value corresponding to the candidate visual media are both false, the electronic device can determine that the matching value corresponding to the candidate visual media is false.
[0047] When the fifth matching value and the sixth matching value corresponding to the candidate visual media are true, the electronic device can determine that the matching value corresponding to the candidate visual media is true.
[0048] In a possible design, the above first target matching value may include a first matching value. The image information of the candidate visual media includes the image feature vector of the candidate visual media, and the information of the search statement includes the text feature vector of the search statement.
[0049] Correspondingly, the process of determining the above first matching value may include:
[0050] The electronic device may input the image feature vectors of each candidate visual media and the text feature vector of the search statement into a binary classification model to obtain the first matching value corresponding to each candidate visual media.
[0051] Among them, the binary classification model is used to determine the matching degree between the candidate visual media and the search statement based on the image feature vector of the candidate visual media and the text feature vector of the search statement, and determine the first matching value corresponding to the candidate visual media according to the matching degree and a preset classification threshold. That is, when the matching degree between the candidate visual media and the search statement is greater than the preset classification threshold, it is determined that the first matching value corresponding to the candidate visual media is true.
[0052] When the matching value between the candidate visual media and the search statement is less than or equal to the preset classification threshold, it is determined that the first matching value corresponding to the candidate visual media is false.
[0053] In this application, the electronic device can use the binary classification model to determine whether the matching degree between the candidate visual media and the search statement is high, so as to determine whether the candidate visual media matches the search statement, realize the secondary determination of the candidate visual media, and ensure the accuracy of the search results.
[0054] Among them, optionally, in the case of the branch corresponding to the base model of the search statement, the image information of the candidate visual media may include the image information of the candidate visual media determined by the base model, and the information of the search statement may include the information of the search statement determined by the base model. Therefore, the image feature vector of the candidate visual media used by the above binary classification model includes the image feature vector of the candidate visual media corresponding to the base model, and the text feature vector of the search statement includes the text feature vector of the search statement corresponding to the base model, ensuring the accuracy of the determination of the first matching value corresponding to the candidate visual media.
[0055] In a possible design, the above first target matching value may include a second matching value. The image information of the candidate visual media includes the similarity between the candidate visual media and the search statement. The information of the search statement may include the search statement or the word segmentation result of the search statement, and the word segmentation result includes at least one word segment.
[0056] Correspondingly, the process of determining the above second matching value may include:
[0057] The electronic device can determine a first similarity threshold corresponding to the search statement based on the length of the search statement. The length of the search statement can be determined by counting the length of the search statement or the word segmentation of the search statement.
[0058] After that, for each candidate visual medium, the electronic device can determine whether the similarity between the candidate visual medium and the search statement is greater than or equal to the first similarity threshold.
[0059] When the similarity between the candidate visual medium and the search statement is greater than or equal to the first similarity threshold, it is determined that the second matching value corresponding to the candidate visual medium is true;
[0060] When the similarity between the candidate visual medium and the search statement is less than the first similarity threshold, it is determined that the second matching value corresponding to the candidate visual medium is false.
[0061] Among them, the first similarity threshold corresponding to the above search statement is positively correlated with the length of the search statement.
[0062] In this application, considering that the longer the search statement is, the more visual content in the visual medium needs to match the search statement. For example, if the search statement includes 3 semantic entities, the more visual content in the visual medium that matches these 3 semantic entities, the more it indicates that the visual medium matches the search statement. Therefore, the electronic device can use the length of the search statement to determine the first similarity threshold corresponding to the search statement. After that, the electronic device can compare the size relationship between the similarity between the candidate visual medium and the search statement and the first similarity threshold to determine that the visual content in the candidate visual medium has a relatively high matching degree with the search statement, so as to obtain the second matching value corresponding to the candidate visual medium, ensuring the accuracy of the determination of the second matching value, and thus can accurately perform a secondary confirmation on the candidate visual medium.
[0063] Among them, optionally, the similarity between the candidate visual medium and the search statement used to determine the second matching value includes the similarity between the candidate visual medium corresponding to the base model and the search statement.
[0064] In a possible design, the above first target matching value may include a third matching value. The image information of the candidate visual medium includes the first label of the candidate visual medium; the information of the above search statement includes the word segmentation of the search statement and the second label;
[0065] Correspondingly, the determination process of the above third matching value may include:
[0066] When the first label of the candidate visual medium includes the word segmentation of the search statement, or includes the second label of the search statement, it is determined that the third matching value corresponding to the candidate visual medium is true;
[0067] When the first tag of the candidate visual media does not include the word segmentation of the search statement and does not include the second tag of the search statement, it is determined that the third matching value corresponding to the candidate visual media is false.
[0068] Among them, when the first tag of the above candidate visual media includes at least one word segmentation of the search statement, the electronic device can determine that the first tag of the candidate visual media includes the word segmentation of the search statement. When the first tag of the above candidate visual media does not include each word segmentation of the search statement, the electronic device can determine that the first tag of the candidate visual media does not include the word segmentation of the search statement. Similarly, when the first tag of the above candidate visual media includes at least one second tag included in the search statement, the electronic device can determine that the first tag of the candidate visual media includes the second tag of the search statement. When the first tag of the above candidate visual media does not include each second tag of the search statement, the electronic device can determine that the first tag of the candidate visual media does not include the second tag of the search statement.
[0069] Alternatively, in order to improve the search accuracy, when the first tag of the above candidate visual media includes each word segmentation of the search statement, the electronic device can determine that the first tag of the candidate visual media includes the word segmentation of the search statement. When the first tag of the above candidate visual media does not include at least one word segmentation of the search statement, the electronic device can determine that the first tag of the candidate visual media does not include the word segmentation of the search statement. Similarly, when the first tag of the above candidate visual media includes each second tag included in the search statement, the electronic device can determine that the first tag of the candidate visual media includes the second tag of the search statement. When the first tag of the above candidate visual media does not include at least one second tag of the search statement, the electronic device can determine that the first tag of the candidate visual media does not include the second tag of the search statement.
[0070] In this application, the electronic device can determine whether the search statement matches the visual content of the candidate visual media by determining whether the search statement hits the first tag of the candidate visual media, so as to realize the secondary determination of the candidate visual media.
[0071] In a possible design, the above first target matching value may include a third matching value. The image information of the candidate visual media includes the first tag of the candidate visual media; the information of the above search statement includes the word segmentation of the search statement.
[0072] Correspondingly, the determination process of the above third matching value may include:
[0073] When the first tag of the candidate visual media includes the word segmentation of the search statement, it is determined that the third matching value corresponding to the candidate visual media is true; when the first tag of the candidate visual media does not include the word segmentation of the search statement, it is determined that the third matching value corresponding to the candidate visual media is false.
[0074] In a possible design, the above first target matching value may include a third matching value. The image information of the candidate visual media includes the first tag of the candidate visual media; the information of the above search statement includes the second tag of the search statement;
[0075] Correspondingly, the process of determining the above third matching value may include:
[0076] When the first tag of the candidate visual media includes the second tag of the search statement, it is determined that the third matching value corresponding to the candidate visual media is true;
[0077] When the first tag of the candidate visual media does not include the second tag of the search statement, it is determined that the third matching value corresponding to the candidate visual media is false.
[0078] In a possible design, the above first target matching value may include a fourth matching value. The image information of the above candidate visual media includes the similarity between the candidate visual media and the search statement, and the information of the search statement includes the search statement;
[0079] The process of determining the above fourth matching value may include:
[0080] When the second preset word list includes the search statement, for each candidate visual media, if the similarity between the candidate visual media and the search statement is greater than or equal to the second similarity threshold, the electronic device may determine that the fourth matching value corresponding to the candidate visual media is true; the second similarity threshold refers to the similarity threshold corresponding to the search statement in the second preset word list;
[0081] If the similarity between the candidate visual media and the search statement is less than the second similarity threshold, the electronic device may determine that the fourth matching value corresponding to the candidate visual media is false.
[0082] In this application, the electronic device may determine whether the search statement hits the second preset word list. When hitting the second preset word list, the electronic device may determine the second similarity threshold corresponding to the search statement from the second preset word list to accurately determine the similarity threshold. After that, the electronic device may compare the size relationship between the similarity between the candidate visual media and the search statement and the second similarity threshold to determine whether the similarity between the candidate visual media and the search statement is large enough to accurately reconfirm the candidate visual media.
[0083] Optionally, the similarity between the candidate visual media used by the electronic device to determine the fourth matching value and the search statement may be the similarity between the candidate visual media corresponding to the base model and the search statement.
[0084] In a possible design, the second target matching value may include a fifth matching value. The image information of the candidate visual media includes the first label of the candidate visual media; the information of the search statement includes the word segmentation of the search statement and the second label.
[0085] Correspondingly, the process of determining the fifth matching value may include:
[0086] When the first label of the candidate visual media includes the word segmentation of the search statement or includes the second label of the search statement, it is determined that the fifth matching value corresponding to the candidate visual media is true.
[0087] When the first label of the candidate visual media does not include the word segmentation of the search statement and does not include the second label of the search statement, it is determined that the fifth matching value corresponding to the candidate visual media is false.
[0088] In a possible design, the second target matching value may include a sixth matching value. The image information of the candidate visual media includes the similarity between the candidate visual media and the search statement, and the information of the search statement includes the search statement.
[0089] Correspondingly, the process of determining the sixth matching value may include:
[0090] When the third preset word list includes the search statement, if the similarity between the candidate visual media and the search statement is greater than or equal to the third similarity threshold, it is determined that the sixth matching value corresponding to the candidate visual media is true; the third similarity threshold refers to the similarity threshold corresponding to the search statement in the third preset word list.
[0091] If the similarity between the candidate visual media and the search statement is less than the third similarity threshold, it is determined that the sixth matching value corresponding to the candidate visual media is false.
[0092] Optionally, the similarity between the candidate visual media used by the electronic device to determine the sixth matching value and the search statement refers to the similarity between the candidate visual media corresponding to the fine-tuning model and the search statement, and the text feature vector of the search statement refers to the text feature vector of the search statement corresponding to the fine-tuning model.
[0093] In a possible design, after obtaining the search statement, the electronic device can filter the search statement to filter out the non-visual semantic entities in the search statement, obtaining a filtered search statement. Subsequently, the electronic device can determine the matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the filtered search statement;
[0094] Determine the first visual media according to the matching value corresponding to the candidate visual media. Moreover, the electronic device can determine the second visual media that matches the non-visual semantic entity in the search statement.
[0095] Subsequently, the electronic device can use the intersection between the first visual media and the second visual media as the target visual media to accurately determine the search result.
[0096] In a possible design, in response to an operation of opening the gallery application, display the first interface;
[0097] Or, in response to an operation of opening the negative first screen triggered on the home screen of the electronic device, display the first interface;
[0098] Or, in response to a pull-down search operation triggered on the home screen of the electronic device, display the first interface.
[0099] In a second aspect, the present application provides an electronic device, which includes a display screen, a memory, and one or more processors; the display screen, the memory, and the processor are coupled; the display screen is used to display the images generated by the processor, the memory is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the electronic device is caused to execute the method as described above.
[0100] In a third aspect, the present application provides a computer storage medium, including computer instructions, which when running on an electronic device, cause the electronic device to execute the method as described above.
[0101] In a fourth aspect, the present application provides a computer program product, which when running on an electronic device, causes the electronic device to execute the method as described above.
[0102] It can be understood that the beneficial effects that can be achieved by the electronic device described in the second aspect, the computer storage medium described in the third aspect, and the computer program product described in the fourth aspect provided above can refer to the beneficial effects in the first aspect and any of its possible design manners, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0103] Figure 1ASchematic diagram one of an interface for visual media search provided by an embodiment of the present application;
[0104] Figure 1B Schematic diagram two of an interface for visual media search provided by an embodiment of the present application;
[0105] Figure 1C Schematic diagram three of an interface for visual media search provided by an embodiment of the present application;
[0106] Figure 1D Schematic diagram of an interface for visual media search provided by an embodiment of the present application Figure Four ;
[0107] Figure 2A Block diagram of an electronic device provided by an embodiment of the present application;
[0108] Figure 2B Software structure diagram of an electronic device provided by an embodiment of the present application;
[0109] Figure 3A Schematic diagram five of an interface for visual media search provided by an embodiment of the present application;
[0110] Figure 3B Schematic diagram of an interface for visual media search provided by an embodiment of the present application Figure Six ;
[0111] Figure 4 Schematic diagram one of a visual media search method provided by an embodiment of the present application;
[0112] Figure 5A Schematic diagram one of the determination process of an image feature vector provided by an embodiment of the present application;
[0113] Figure 5B Schematic diagram two of the determination process of an image feature vector provided by an embodiment of the present application;
[0114] Figure 5C Schematic diagram of the determination process of a text feature vector provided by an embodiment of the present application;
[0115] Figure 5D Schematic diagram of the determination process of a similarity provided by an embodiment of the present application;
[0116] Figure 6 Schematic diagram two of a visual media search method provided by an embodiment of the present application;
[0117] Figure 7 Schematic diagram three of a visual media search method provided by an embodiment of the present application;
[0118] Figure 8Schematic diagram of a visual media search method provided by an embodiment of the present application Figure Four ;
[0119] Figure 9A Schematic diagram one of a secondary confirmation provided by an embodiment of the present application;
[0120] Figure 9B Schematic diagram one of a visual media provided by an embodiment of the present application;
[0121] Figure 10 Schematic diagram five of a visual media search method provided by an embodiment of the present application;
[0122] Figure 11A Schematic diagram one of a matrix provided by an embodiment of the present application;
[0123] Figure 11B Schematic diagram two of a visual media provided by an embodiment of the present application;
[0124] Figure 11C Schematic diagram two of a matrix provided by an embodiment of the present application;
[0125] Figure 11D Schematic diagram of a positive and negative sample discrimination process provided by an embodiment of the present application;
[0126] Figure 11E Schematic diagram of a binary classification model training process provided by an embodiment of the present application;
[0127] Figure 12 Schematic diagram two of a secondary confirmation provided by an embodiment of the present application;
[0128] Figure 13 Schematic diagram three of a secondary confirmation provided by an embodiment of the present application;
[0129] Figure 14A Schematic diagram of a secondary confirmation provided by an embodiment of the present application Figure Four ;
[0130] Figure 14B Schematic diagram five of a secondary confirmation provided by an embodiment of the present application. Detailed implementation manners
[0131] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this embodiment, unless otherwise specified, the meaning of "a plurality" is two or more.
[0132] To better understand the embodiments of the present application, the following first explains the vocabulary involved in the present application.
[0133] Visual media: refers to pictures or videos.
[0134] Semantic entity: The named entity recognition (NER) technology can identify sentences and recognize entities with specific meanings in the text, such as people's names and place names. In this solution, the entities with specific meanings identified are called semantic entities.
[0135] Visual content-related and visual content-unrelated: Visual content refers to the objects presented by visual media and their interrelationships, etc. Simply put, visual content can be understood as the content included in visual media. In the context of image search in the present application, the data that can be obtained after the natural picture understanding of visual media files by the model is called "visual content-related". This solution refers to the data that is related to visual media files and can be obtained without the picture understanding ability of the model as "visual content-unrelated". For example, when an electronic device collects visual media files, it can obtain and save the shooting location, shooting time, name, file attributes, etc.
[0136] For example, in the sentence "The photo taken in City 1 this year", "this year" (shooting time), "City 1" (shooting location), and "photo" (file attribute) are all data that can be obtained and saved by the electronic device when collecting visual media files. Therefore, "this year", "City 1", and "photo" are irrelevant to visual semantics; in the sentence "The sky taken in City 1 this year", "sky" can only be obtained through the picture understanding ability of the model to understand the picture. Therefore, "sky" is related to visual semantics.
[0137] Text semantic vector: It can be obtained by sending text into a text encoder. It is a vector that can represent the semantic features of the entire sentence. The text encoder can use a clip model or other models, such as the Transformer commonly used in natural language processing (NLP). This solution does not limit this here. Among them, in the present application, the text semantic vector can also be called the text feature vector.
[0138] Visual semantic vector: It can be obtained by sending visual media (such as images) into an image encoder. The image encoder can use a clip model or other models, such as a CNN model or a VIT model. This solution does not limit this here. Among them, in the present application, the visual semantic vector can also be called the image feature vector.
[0139] Vector similarity: used to describe the similarity between two vectors (e.g., between a text semantic vector and a visual semantic vector). In the embodiments of the present application, the visual media matching the search statement can be determined by comparing the similarity between the text semantic vector of the search statement and the visual semantic vector of the visual media. Generally, the vector similarity can be calculated through the cosine similarity calculation formula. Of course, it can also be calculated by other means.
[0140] An electronic device (such as a mobile phone) can manage visual media such as pictures and videos of users through a gallery application. Taking the example of a mobile phone taking a photo, after the mobile phone takes a photo, the gallery application determines the attribute tags such as the shooting location, shooting time, and photo name corresponding to the photo, and can use the attribute tags as the index of the picture. After the gallery application establishes an index for the visual media, it can provide corresponding search services to the user. Specifically, the user can search for pictures or videos on the mobile phone by entering keywords in the gallery application. Exemplarily, the user can enter keywords such as "sky", "cat", "time point 1" in the search box provided by the gallery application, and the gallery application matches the keywords entered by the user with the indexes of visual media such as pictures and videos in the gallery application to obtain search results.
[0141] Optionally, the above attribute tags may further include attributes such as the face identity identification number (industrial design, ID), name of the person in the photo, and the relationship between the person and the mobile phone user. Among them, the face ID of the person in the photo can be automatically generated by the gallery application, and the name of the person in the photo and the relationship between the person and the mobile phone user can be manually input by the user. In practical applications, the same person corresponds to the same face ID, the same name, and the same relationship with the mobile phone user. Therefore, in order to simplify the user operation, the user only needs to input the name of the person and the relationship with himself once for the same person. Subsequently, the gallery application automatically configures the face ID, name of the person, and the relationship between the person and the mobile phone user for the pictures containing the face of the person through face recognition technology. In addition, the above attribute information such as the photo name can also be automatically generated by the gallery application or manually named by the user.
[0142] The following is an exemplary description of the interface involved in the search process of the gallery application with reference to the accompanying drawings:
[0143] As Figure 1A shown in (a) of, the mobile phone can display the main interface 101, and the main interface can also be called the desktop. The main interface 101 may include the icon 102 of the gallery application. The mobile phone receives the operation of the user clicking the icon 102, and in response to this operation, the mobile phone can start the gallery application and display as Figure 1AThe interface 103 shown in (b) therein, where the interface 103 can be an album interface. It should be noted that in response to the user's operation of clicking the icon 102, the mobile phone can launch the gallery application and display the photo interface of the gallery. The photo interface includes thumbnails of the photos (i.e., pictures) in the gallery or a large image of a certain photo. In the photo interface, in response to the user's operation on the "Album" control, the above-mentioned album interface 103 is displayed.
[0144] As Figure 1A shown in (b) therein, the interface 103 includes multiple albums. Among them, the "All Photos" album includes 2,023 photos, the "Camera" album includes 1,502 photos and videos, the "Screenshots & Screen Records" album includes 102 photos and videos, the "My Favorites" album includes 48 photos and videos, the "Multiple Gains from One Recording" album has 34 photos and videos, the "Video Editing" album has 65 videos, the "Self-created Album" has 57 photos and videos, and the "Shared Album" has 100 photos and videos.
[0145] As Figure 1A shown in (b) therein, the interface 103 may include a search box 104. The mobile phone can receive the user's operation of clicking the search box 104. In response to this operation, the mobile phone can display the interface 105 shown in (c) as Figure 1A therein. This interface 105 can be called a search interface. Among them, the interface 105 can display the classification information of the photos to the user. For example, in the interface 105, the mobile phone classifies the photos of the local machine according to time, people, and things, etc. For example, in the dimension of time, the mobile phone classifies the photos of the local machine according to three time periods: "This Month", "Last Month", and "This Year". Among them, the "This Month" album includes the photos or videos taken by the mobile phone this month, the "Last Month" album includes the photos or videos taken by the mobile phone last month, and the "This Year" album includes the photos or videos taken by the mobile phone this year. In the dimension of people, the mobile phone classifies the photos of the local machine according to different people, such as the four different people in the interface 105. In the dimension of things, the mobile phone classifies and displays the photos of the local machine according to "Scenery", "Animals", "Documents", and "Buildings". It should be noted that the above classification dimensions can also be others, and no specific restrictions are made here. In the interface 105, the user can see this classification information without entering keywords.
[0146] Optionally, the interface 105 may further include a search history 107 and an option to "clear" 108. The search history includes keywords that the user has entered, such as "flowers", "coffee", "cats", etc. The mobile phone can receive the operation of the user clicking on "clear" 108, and in response to this operation, the mobile phone can clear the search history. After the mobile phone clears the search history, the keywords that the user has entered are no longer displayed on the search interface 105. For example, in response to the operation of the user clicking on "clear" 108, as Figure 1B shown, the search history 107 and the option to "clear" 108 are no longer displayed on the search interface 105, and the content displayed below moves up.
[0147] In response to the operation of the user entering the keyword "sky" on the interface 105, the mobile phone displays the interface 109 as shown in Figure 1C (a). Among them, the mobile phone can search for data on the local machine related to the keyword "sky". Specifically, the mobile phone can associate the keyword "sky" to obtain associated words such as "sky" and photos containing the word "sky". Then, search according to each associated word to obtain the search results of each associated word. For example: 100 photos related to "sky" and 32 photos related to photos containing the word "sky". Among them, the 100 photos related to "sky" can be recalled because the classification labels of each of these 100 photos match "sky" or "sky" and its associated words; the 32 photos related to photos containing the word "sky" can be recalled because through optical character recognition (OCR) technology, it is recognized that these 32 photos contain the character "sky". The union of the search results of each of these multiple associated words can be used as the search result of the keyword "sky".
[0148] The interface 109 also displays some search results of the keyword "sky" and a "more" option 110 corresponding to the search results of the keyword "sky". The mobile phone receives the click operation of the user on the "more" option 110 and displays the interface 111 as shown in Figure 1C (b). Among them, the interface 111 is used to display photos and videos in the search results of the keyword "sky". Optionally, the photos and videos can be classified and displayed according to time. In addition, the interface 111 also includes a return key 112 and a title 113. In response to the operation of the user on the return key 112, the mobile phone can redisplay the interface 109. The title 113 may include the keyword "sky".
[0149] That is to say, in the gallery application, when the user enters a simple search statement in the search box, such as a simple keyword, for example: sky, location 1, time 1, etc., corresponding search results can be obtained. However, because the mobile phone's ability to understand and associate search statements is limited, when the user enters a more complex search statement in the search box, if the keywords in the search statement cannot match the attribute tags of the pictures or the text in the pictures, no photos may be found. That is to say, the existing gallery applications do not support the search function based on complex search statements. As Figure 1D shown, when the user enters a more complex search statement "warming oneself by the fire and brewing tea" in the search box of interface 114, the mobile phone cannot understand the associated words of "warming oneself by the fire and brewing tea". Since the photos do not have tags that can match "warming oneself by the fire and brewing tea" or its associated words, the search result shows "no pictures".
[0150] In some embodiments, the electronic device can use an encoder to calculate the similarity between the visual media on the electronic device and the search statement, so that the electronic device can use the visual media with a similarity higher than the threshold as the search result. Then, the electronic device displays the search result to show the visual media required by the user. For example, the encoder can include a text encoder and an image encoder. The electronic device can input each visual media on the electronic device into the image encoder to obtain the image feature vector corresponding to each visual media. And the electronic device can input the search statement into the text encoder to obtain the text feature adjacent corresponding to the search statement. Then, for each visual media, the electronic device can calculate the similarity between the visual media and the search statement based on the image feature vector corresponding to the visual media and the text feature vector corresponding to the search statement.
[0151] For another example, the above encoder can be an encoder. The electronic device can input each visual medium and search statement on the electronic device into the encoder respectively, so that the encoder can obtain the similarity between each visual medium and the search statement according to the image feature vector of each visual medium and the text feature vector of the search statement. However, the electronic device may not be able to accurately determine the image feature vector of the visual medium only by using the encoder, resulting in a low calculation accuracy of the similarity between the visual medium and the search statement, so that the search results displayed by the electronic device may not be the visual media required by the user and cannot meet the search needs of the user. For example, when the search statement includes landmark buildings, the encoder may not be able to accurately determine the image feature vector of the visual medium including landmark buildings, resulting in a low accuracy of the calculated similarity between the visual medium and the search statement, and further causing the electronic device to be unable to accurately search for the visual media required by the user. Simply put, the electronic device cannot accurately identify the visual media including landmark buildings, resulting in a low accuracy of the determined search results. For another example, when the search statement includes rare animals, the encoder may not be able to accurately identify the visual media including non-scene animals, resulting in a low accuracy of the determined search results.
[0152] Therefore, in order to improve the accuracy of visual media search to meet the search needs of users, the present application provides a visual media search method. After the electronic device determines the visual media with a similarity higher than the threshold by using a relevant model (such as a multimodal model), the visual media with a similarity higher than the threshold is used as a candidate visual media. Then, the electronic device can continue to perform a secondary confirmation on the candidate visual media to determine whether the candidate visual media is the visual media required by the user, that is, to obtain the target visual media, so as to accurately determine the retrieval result. Then, the electronic device can display the target visual media, that is, display the visual media required by the user, to meet the search needs of the user.
[0153] Exemplarily, the above electronic device can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., which can store visual media.
[0154] Exemplarily, Figure 2AThe structural schematic diagram of the electronic device 200 is shown. The electronic device 100 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone interface 270D, a sensor module 280, a key 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.
[0155] Among them, the sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.
[0156] It can be understood that the structure schematically shown in the embodiments of this application does not constitute a specific limitation on the electronic device 200. In other embodiments of this application, the electronic device 200 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0157] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0158] The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0159] A memory can also be provided in the processor 210 for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory can store the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.
[0160] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0161] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative and do not constitute a structural limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may also adopt different interface connection methods or a combination of multiple interface connection methods in the above embodiments.
[0162] The electronic device 200 realizes the display function through a GPU (Graphics Processing Unit), a display screen 294, and an application processor, etc. The GPU is connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change display information.
[0163] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 294, where N is a positive integer greater than 1.
[0164] The electronic device 200 can implement the shooting function through an ISP, a camera 293, a video codec, a GPU, a display screen 294, an application processor, etc.
[0165] The camera 293 is used to capture static images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB, YUV, etc. In some embodiments, the electronic device 100 may include one or N cameras 293, where N is a positive integer greater than 1.
[0166] The video codec is used to compress or decompress digital videos. The electronic device 200 can support one or more video codecs. In this way, the electronic device 200 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0167] The NPU is a computing processor for neural networks (NN). By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously learn on its own. Through the NPU, applications such as intelligent cognition of the electronic device 200 can be realized, such as image recognition, face recognition, speech recognition, text understanding, etc.
[0168] The external memory interface 220 can be used to connect to an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 200. The external memory card communicates with the processor 210 through the external memory interface 220 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0169] The internal memory 221 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 221 can include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.). The data storage area can store data created during the use of the electronic device 200 (such as audio data, phone book, etc.). In addition, the internal memory 221 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. The processor 210 executes various functional applications and data processing of the electronic device 200 by running the instructions stored in the internal memory 221 and / or the instructions stored in the memory provided in the processor.
[0170] Figure 2B It is a software structure block diagram of the electronic device 200 in the embodiments of the present application. The software system of the electronic device 200 can adopt a layered architecture. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. Taking the Android system as an example, in some embodiments, the Android system can be divided into four layers, from top to bottom are the application layer, application framework layer, Android runtime and system libraries, and kernel layer.
[0171] As Figure 2B shown, the application layer can include application programs such as a gallery service module, a search module, a multimodal understanding module, a natural language understanding module, and a camera application.
[0172] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.
[0173] The system library can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (e.g., OpenGL ES), 2D graphics engine (e.g., SGL), etc. Among them, the media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0174] The kernel layer is the layer between hardware and software.
[0175] Next, in combination with the capture and photo-taking scenario, the working processes of the software and hardware of the electronic device 200 will be exemplarily described.
[0176] When the touch sensor 280K receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including information such as touch coordinates and the timestamp of the touch operation). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking the touch operation as a touch click operation and the control corresponding to the click operation being the control of the camera application icon as an example, the camera application calls the interface of the application framework layer to start the camera application, and then starts the camera driver by calling the kernel layer to capture a static image or video through the camera 293.
[0177] Taking the electronic device as a mobile phone as an example, the visual media search method provided by the embodiments of the present application will be introduced. The visual media search method provided by the embodiments of the present application can be applied to application programs such as the gallery application and the file management application.
[0178] Next, the interfaces and search logics related to the visual media search method provided by the embodiments of the present application will be exemplarily described in combination with the accompanying drawings.
[0179] As Figure 3A shown in (a) of, the interface 301 displays search history 303 and a "Clear" option 304. The search history 303 includes search statements that the user has entered, such as: "Watching the sunrise on the top of the mountain", "The sky taken in the photo". Other contents displayed on the interface 301 can refer to the relevant display contents of the above interface 105, which will not be elaborated here. Among them, the interface 301 can be called a search interface. The mobile phone can respond to the user'sFigure 1A Clicking on the search box 104 in the interface 103 shown in (b) in
[0180] As Figure 3A shown in (b) in Figure 3A the interface 305 (i.e., the search result interface), the user enters the search statement "warming the stove and brewing tea" in the search box 306 of the interface 305, and the mobile phone searches for 239 pictures. The mobile phone displays some search results of the search statement "warming the stove and brewing tea" (e.g., thumbnails of 8 pictures) and the "More" option 307 corresponding to the search results of the search statement "warming the stove and brewing tea" on the interface 305. In response to the user's operation on the "More" option 307, the mobile phone displays the interface 308 shown in (c) in
[0181] In addition, the mobile phone can also be provided with a negative first screen, a pull-down search interface, etc. It can be understood that the negative first screen can be the leftmost split screen of the electronic device, which is used to provide functions such as search and quick services for the user. Among them, the negative first screen can also be used to display notification messages to be pushed to the user, such as application messages subscribed by the user, real-time hot search messages, segment selection, itinerary information, etc. The pull-down search interface is an interface displayed in response to the user's pull-down operation on the main interface, and this interface is used to provide functions such as search and application suggestions for the user, and this interface and Figure 3B the interface 315 in (c) below can be the same interface.
[0182] Next, an example will be given with the negative first screen. When the user needs to view the negative first screen of the mobile phone, the user can slide the screen of the mobile phone to make the electronic device display the negative first screen.
[0183] Exemplarily, as shown in (a) in Figure 3B the mobile phone can receive operation 1 implemented by the user on the interface 309 (which can be called the desktop) of the mobile phone. Exemplarily, this operation 1 can be a rightward sliding operation as shown in (a) in Figure 3B In response to this operation 1, the mobile phone can display the negative first screen 310 shown in (b) in Figure 3B Among them, the negative first screen 310 can include: a search box 311, quick services 312, default cards 313, recommended cards 314, etc. The quick services 312 can be quick access to a certain page or function of an application program, such as: scan code, payment code, ride code, etc.; the default cards can be: gallery cards, remaining battery cards, etc.; the recommended cards can be recommended application cards.
[0184] The mobile phone receives a click operation by the user on the search box 311 on the minus one screen 310. In response to this click operation, the mobile phone can display the interface 315 as shown in Figure 3B figure (c). The interface 315 may include: a search box 316 and application suggestions. The application suggestions include: icons of applications recommended for use. The interface 315 may also include: a search history 317 and its corresponding "clear" option 318. In response to a triggering operation by the user on the "clear" option 318, the search history 317 and the "clear" option 318 are no longer displayed on the interface 315. In addition, a hot news title, such as "Marathon race", may be displayed in the search box 316.
[0185] As Figure 3B shown in figure (d), in the interface 319, a search statement "mountains photographed on the weekend" input by the user is displayed in the search box. A preview area 322 of the search results corresponding to the search statement "mountains photographed on the weekend" and a "search in the application" option 323 corresponding to the gallery application are also displayed in the interface 319. In response to a triggering operation by the user on the preview area 322, the mobile phone can enter a photo details interface provided by the gallery application for the user to flip through the search results corresponding to the search statement "mountains photographed on the weekend", and visual media matching the search statement, such as photos or videos. In response to a triggering operation by the user on the "search in the application" option 323, the mobile phone displays the interface 324 provided by the gallery application as shown in Figure 3B figure (e). The interface 324 displays some search results of the search statement "mountains photographed on the weekend" and a "more" option corresponding to the search results of the search statement "mountains photographed on the weekend". In response to a triggering operation by the user on this "more" option, the mobile phone can display a search result details interface, and the search result details interface shows the pictures in the search results of the search statement "mountains photographed on the weekend". Additionally, an online search option 321 may also be displayed in the interface 319. In response to a triggering operation by the user on the online search option 321, the mobile phone displays a search web page and shows online search results in the search web page.
[0186] The interfaces and search logics involved in the visual media search method are introduced above. Next, the specific implementation process of the visual media search method will be continued in combination with the Figure 2B software structure shown above. As Figure 4 shown, this implementation process may include S401 - S423. Among them, S401 - S407 may belong to the index construction stage, and S408 - S423 may belong to the search stage.
[0187] S401. The gallery service module adds and / or modifies visual media and their attributes.
[0188] The attributes of the visual media may include one or more of the collection location, collection time, visual media name, face ID of the person in the visual media, person name, and the relationship between the person and the mobile phone user. Taking a captured video or picture as an example, the collection location refers to the shooting location, and the collection time refers to the shooting time; taking a screenshot as an example, the collection location refers to the screenshot location, and the collection time refers to the screenshot time; taking a downloaded video or picture as an example, the collection location refers to the download location, and the collection time refers to the download time.
[0189] The user can add visual media by means such as shooting, downloading, and taking screenshots. In addition, the user can also modify the existing visual media to obtain new visual media. Such modifications include but are not limited to operations such as beautification, custom naming, and adding watermarks.
[0190] S402. The gallery service module stores the visual media and its attributes.
[0191] The gallery service module can, in response to the addition or modification operation of the visual media input by the user, store the visual media and its attributes locally on the mobile phone. In practical applications, with the authorization of the user, the mobile phone can store the locally stored visual media and its attributes in the cloud to relieve the storage pressure on the local mobile phone.
[0192] S403. The gallery service module sends Request 1 to the multimodal understanding module. Among them, Request 1 is used to trigger the visual semantic understanding of the visual media.
[0193] S404. The multimodal understanding module returns the image feature vector of the visual media to the gallery service module.
[0194] In the embodiment of the present application, the multimodal understanding module, in response to the above Request 1, determines the image feature vector of the visual media stored in the mobile phone. Since visual semantic understanding requires a large amount of computing resources, in order not to affect the user's use, when the mobile phone is in the charging and screen-off state, the gallery service module can request the multimodal understanding module to perform visual semantic understanding on the visual media stored locally on the mobile phone, such as performing visual semantic understanding on the newly added or modified visual media to obtain the visual semantic vector (or called the image feature vector) of the visual media, so as to realize the offline processing of the visual media and reduce the impact of the visual semantic understanding of the visual media on other services running on the mobile phone.
[0195] Among them, the multimodal understanding module can perform visual semantic understanding on visual media based on a multimodal model to obtain an image feature vector of the visual media. In addition, the multimodal model can be used not only for: performing visual semantic understanding on visual media to obtain a visual semantic vector of the visual media; but also for: performing semantic understanding on a search statement to obtain a text feature vector (or called a sentence semantic vector) of the search statement. The process of the multimodal model performing semantic understanding on the search statement can refer to the relevant descriptions below and will not be introduced in detail here.
[0196] In some embodiments, the multimodal model can map visual media and text into vectors of the same dimension. That is to say, the dimension of the visual semantic vector of visual media is the same as the dimension of the semantic vector of text (for example: the sentence semantic vector of a search statement). The multimodal model can specifically be a contrastive language-image pre-training (CLIP) model. The mobile phone can map visual media and text into a unified vector space through the CLIP model to understand the relationship between different modal resources visually and textually, and then use it for image retrieval.
[0197] Among them, the CLIP model is a standard CLIP model, that is, an existing CLIP model, or the CLIP model is a custom CLIP model.
[0198] Exemplarily, the above custom CLIP model can include a base model and a fine-tuning model. Among them, the fine-tuning model is obtained by continuing to train the base model with a training set of a specific scenario on the basis of the base model. Therefore, the fine-tuning model has a supported range and can support the processing of images and their texts in a specific scenario. The generalization ability of the fine-tuning model is less than that of the base model, while the retrieval ability of the base model in a specific scenario is less than that of the fine-tuning model.
[0199] In some embodiments, the above-mentioned base model and fine-tuning model reuse part of the network to reduce the resource occupancy of the mobile phone installed with the custom CLIP model. The base model and the fine-tuning model need to process not only visual media but also text. The module for processing visual media in the base model can be called the visual encoding module, and the module for processing visual media in the fine-tuning model can be called the visual fine-tuning module. The above-mentioned image feature vectors can include the image feature vectors corresponding to the base model and the image feature vectors corresponding to the fine-tuning model. The visual encoding module of the base model and the visual fine-tuning module of the fine-tuning model can reuse the image encoder, so that the visual encoding module can use the feature vectors of the visual media output by the image encoder to continue to determine the first L2 norm of the visual media, realizing the determination of the image feature vectors corresponding to the base model, and the visual fine-tuning module can use the feature vectors of the visual media output by the image encoder to continue to determine the second L2 norm of the visual media, realizing the determination of the image feature vectors corresponding to the fine-tuning model.
[0200] Exemplarily, after receiving Request 1 sent by the gallery service module, in response to this Request 1, as Figure 5A shown, the multimodal understanding module can input k visual media in the mobile phone into the visual encoding module. The image encoder in the visual encoding module performs encoding processing on each of the k visual media to obtain the image feature vectors 1 of each visual media. The image feature vector of this visual media can be 768-dimensional, that is, X = {x1, x2,..., x768}, and X represents the image feature vector 1. After that, on the one hand, the visual encoding module can output the 768-dimensional image feature vector 1 so that the image feature vector 1 continues to be used as the input parameter of the visual fine-tuning module. On the other hand, the visual encoding module continues to use the mapping matrix 1 and the image feature vector 1 to calculate and output the first L2 norm α1 of the image feature vectors of each of the k visual media. Among them, the dimension of the mapping matrix 1 is 768*512-dimensional to map the image feature vector 1 from 768 dimensions to 512 dimensions. Here, the image feature vector 1 of the visual media and the first L2 norm of the visual media can be regarded as the image feature vectors of the visual media corresponding to the base model.
[0201] Specifically, for each of the k visual media, the visual encoding module can adopt to calculate the first L2 norm α1 of the image feature vector of the visual media. Among them, X represents the 768-dimensional image feature vector 1, M 1 represents the mapping matrix 1, V 1 represents the image feature vector 2, and the dimension of V 1 is 512-dimensional, represents the i-th element in V 1 .
[0202] As Figure 5BAs shown in the figure, after the visual fine-tuning module receives the image feature vector 1 of each of the k visual media, for each visual media, it continues to use the mapping matrix 2 and the image feature vector 1 of this visual media to calculate and output the second L2 norm α2 of the image feature vector of this visual media. Among them, the dimension of the mapping matrix 2 is 768*512, so as to map the image feature vector 1 from 768 dimensions to 512 dimensions. Here, the image feature vector 1 of the visual media and the second L2 norm of the visual media can be regarded as the image feature vector of the visual media corresponding to the fine-tuning model.
[0203] Specifically, for each of the k visual media, the visual encoding module can adopt to calculate the second L2 norm α2 of the image feature vector of the visual media. Among them, X represents the 768-dimensional image feature vector 1, and M 2 represents the mapping matrix 2, and V 2 represents the image feature vector 3, and the dimension of V 2 is 512 dimensions, represents the i-th element in V 2 .
[0204] In some embodiments, the above-mentioned fine-tuning model (such as the mapping matrix 2 in the above-mentioned visual fine-tuning module) is obtained by continuing to train using a training set of a specific scenario. When the device trains the fine-tuning model, it can freeze the image encoder and only train the last fully connected layer M 2 of the fine-tuning model. The training set of this specific scenario can be an image training set including visual content corresponding to a preset whitelist. The preset whitelist represents the support range of the fine-tuning model. Among them, the device can be the above-mentioned mobile phone or not, and the present application does not limit the device for training the fine-tuning model.
[0205] After obtaining the image feature vector of the visual media (such as the image feature vector of the visual media corresponding to the above-mentioned base model and the image feature vector of the visual media corresponding to the fine-tuning model), in order to facilitate retrieving the visual media required by the user using the image feature vector of the visual media, the multimedia understanding module can save the image feature vector of the visual media, that is, the above-mentioned image feature vector 1, the first L2 norm α1 and the second L2 norm α2, so that only a 770-dimensional visual vector needs to be saved for saving the image feature vector of a visual media, instead of saving 512*2 dimensions, that is, 1024-dimensional visual features, reducing the resources required to store the image feature vector. Among them, the 512-dimensional vector is the image feature vector obtained by processing the 768-dimensional vector using the mapping matrix 1 or the mapping matrix 2 (that is, the above-mentioned V 1 and V 2 ).
[0206] It should be noted that the above image encoder located in the visual encoding module is only an example. The image encoder can also be located in the visual fine-tuning module, and the present application does not limit it. In addition, the image feature vectors of the visual media mentioned above including the image feature vectors of the visual media corresponding to the base model and the image feature vectors of the visual media corresponding to the fine-tuning model are only an example. The image feature vectors of the visual media can also only include the image feature vectors determined by one model, and this model can be a clip model or not a clip model.
[0207] As mentioned above, the determination of the image feature vectors of the visual media in the mobile phone can be determined in an offline state, while the text feature vectors of the search statements can be determined online after the mobile phone receives the search statements input by the user, so that the mobile phone can use the text feature vectors of the search statements and the image feature vectors of the visual media in the mobile phone to determine the visual media that matches the search statements. Among them, the process of determining the text feature vectors of the search statements and using the text feature vectors of the search statements and the image feature vectors of the visual media to determine the visual media that matches the search statements can be referred to below, and will not be introduced in detail here. Next, the process of establishing the index of the visual media will be introduced first.
[0208] S405. The gallery service module stores the image feature vectors of the visual media.
[0209] S406. The gallery service module sends the attribute information of the visual media and its visual semantic vectors to the search module.
[0210] In the embodiment of the present application, the gallery service module can store the image feature vectors of each of the k visual media locally in the mobile phone. And the gallery service module can send the attribute information of the visual media and its image feature vectors to the search module in batches for the search module to construct the index of the visual media.
[0211] Among them, optionally, the multi-gallery service module may not upload the image feature vectors of the visual media to the cloud. Of course, with the authorization of the user, the image feature vectors of the visual media can also be uploaded to the cloud, and the present application does not limit it.
[0212] S407. The search module constructs the index of the visual media.
[0213] The index of the visual media constructed by the search module may include: the attributes of the visual media and / or the visual semantic vectors of the visual media.
[0214] S408. The gallery service module receives the search statements input by the user.
[0215] The user can pass the search interface provided by the gallery service module. For example, the user can Figure 3AEnter "warming the tea around the stove" in the search box 306 on the interface 305 shown in (b) of []. This "warming the tea around the stove" is the search statement.
[0216] S409. The gallery service module sends the search statement to the search module.
[0217] S410. The search module determines whether the search statement includes visual content.
[0218] In the embodiment of the present application, the search module can determine whether the search statement includes visual content, that is, determine whether it is necessary to search for visual media using the image feature vector and the text feature vector. In other words, the search module determines whether it is necessary to use the multimodal model to determine the visual media.
[0219] If the search statement does not include visual content, it indicates that the visual media can be searched using the attributes of the visual media, and there is no need to search for the visual media using the image feature vector and the text feature vector. That is, it indicates that there is no need to use the multimodal model to determine the visual media, and the search module can execute S411.
[0220] If the search statement includes visual content, it indicates that it is necessary to search for visual media using the image feature vector and the text feature vector. That is, it indicates that it is necessary to use the multimodal model to determine the visual media, and the search module can execute S412.
[0221] In some embodiments, the search module may determine whether a search statement includes visual content by determining whether the search statement includes a visual semantic entity. The search module may send a request 2 to the natural language understanding module. The natural language understanding module may perform semantic entity recognition on the search statement to obtain the semantic entities included in the search statement. Among them, the natural language understanding module performs semantic entity recognition based on a natural language understanding model. Specifically, the natural language understanding module may use named entity recognition (NER) technology to perform semantic entity recognition on the search statement to obtain the semantic entities included in the search statement. In the embodiments of the present application, the semantic entity may also be referred to as an entity, and the semantic entity may include one or more of: a semantic entity related to time, a semantic entity related to location, a semantic entity related to a person's name, a semantic entity related to a person's relationship, and a semantic entity related to visual content (or referred to as a visual semantic entity). Among them, optionally, the semantic entity related to visual content is determined from M (M≥1) preset semantic entities related to visual content through named entity recognition technology. These M semantic entities can be configured by developers of the gallery service module according to actual situations. Generally, these M semantic entities are all nouns. Exemplarily, for the search statement "the sky photographed in City 1 on September 1st", semantic entity recognition can obtain three semantic entities: "September 1st", "City 1", and "the sky". Among them, "the sky" is related to visual content and can be a visual semantic entity.
[0222] After that, the natural language understanding module may return the recognized semantic entities to the search module. After that, the search module may determine whether the semantic entities include a visual semantic entity. In the case where the semantic entities do not include a visual semantic entity, the search module may determine that the search statement does not include visual content. In the case where the semantic entities include a visual semantic entity, the search module may determine that the search statement includes visual content.
[0223] S411. The search module performs recall based on the index of the visual media and the search statement to obtain search results.
[0224] Exemplarily, when the search statement does not include a visual semantic entity, it indicates that the search statement does not include content related to visual semantics. The search module can directly query the visual media corresponding to the semantic entity based on the constructed index, that is, based on the attributes of each visual media, and use it as the search result. Here, the semantic entity represents the attributes of the visual media. For example, if the search statement is "September 1st" and does not include visual content, the search module can, based on the index, find the visual media with the acquisition time of September 1st to obtain the search result. Another example is that the search statement includes "photos of Zhang San" and does not include visual content. The search module can match the names of each visual media with "Zhang San" to determine the visual media with the name attribute of "Zhang San" and use it as the search result.
[0225] S412. The search module filters out the non-visual semantic entities in the above search statement to obtain a filtered search statement.
[0226] Exemplarily, when the above semantic entity includes a visual semantic entity, it indicates that the search statement includes content related to visual semantics and may include content unrelated to visual semantics, that is, non-visual semantic entities. Since non-visual semantic entities are irrelevant to the search for visual content, the search module can first filter out the non-visual semantic entities in the search statement to obtain a filtered search statement for searching for visual media that matches the filtered search statement.
[0227] Among them, optionally, non-visual semantic entities refer to: semantic entities related to time, semantic entities related to location, semantic entities related to personal names, semantic entities related to personal relationships, etc., which are semantic entities related to the attributes of visual media. It should be understood that since personal relationships have been regarded as the attributes of visual media before, personal relationships can be regarded as non-semantic entities here.
[0228] For example, for the search statement "sky photographed this year", "this year" is a non-visual semantic entity. Correspondingly, the filtered search statement is "photographed sky".
[0229] In practical applications, after filtering out the semantic entities irrelevant to visual content, there may be some redundant stop words. For example, for the search statement "sky photographed in City 1 this year", after deleting "this year" and "City 1", the stop word "in" becomes a redundant word, so the search module can also filter it out. Specifically, the search module can filter out the non-visual semantic entities and their related stop words in the search statement to obtain a filtered search statement. For example, the filtered search statement corresponding to the search statement "sky photographed in City 1 this year" is "photographed sky".
[0230] S413. The search module sends the filtered search statement to the multimodal understanding module.
[0231] S414. The multimodal understanding module performs semantic understanding on the filtered search statement to obtain the text feature vector of the filtered search statement.
[0232] Exemplarily, the multimodal understanding module may use a multimodal model to determine the text feature vector of the filtered search statement. Optionally, the multimodal model may be a CLIP model. The CLIP model may be a standard CLIP model, or the CLIP module is a custom CLIP model.
[0233] In some embodiments, from the above, it can be seen that the above image feature vector may include the image feature vector corresponding to the base model and the image feature vector corresponding to the fine-tuning model. The visual encoding module of the base model and the visual fine-tuning module of the fine-tuning model may reuse the image encoder. Correspondingly, in order to keep the similarity of the text-image pair consistent, the present application introduces a text encoding module, so that the text encoding module can reuse the text encoder to respectively output the text feature vector of the filtered search statement corresponding to the base model and the text feature vector of the filtered search statement corresponding to the fine-tuning model, for calculating the similarity between the search statement and the visual media by using the text feature vector of the filtered search statement corresponding to the base model and the image feature vector of the visual media corresponding to the base model, and calculating the similarity between the search statement and the visual media by using the text feature vector of the filtered search statement corresponding to the fine-tuning model and the image feature vector of the visual media corresponding to the fine-tuning model, to implement the calculation of the similarity of the text-image pair.
[0234] Exemplarily, as Figure 5C shown, the multimodal understanding model inputs the filtered search statement into the CLIP model. The text encoder in the text encoding module of the CLIP model encodes the filtered search statement to obtain and output the feature vector of the filtered search statement, and the dimension of this feature vector is 768. Then, the text encoding module may respectively input the feature vector of the filtered search statement into the base model branch and the fine-tuning model branch. Among them, the structures of the base model branch and the fine-tuning model branch are similar.
[0235] For the base model branch: The text encoding module uses the mapping matrix 3 and the feature vector of the filtered search statement to obtain the intermediate variable 1 with a dimension of 512. Exemplarily, the text encoding module calculates the intermediate variable 1 according to 1 T 1 =YN 1 where T 1 represents the intermediate variable 1, Y represents the feature vector of the filtered search statement, and N represents the mapping matrix 3. The text encoding module calculates the first L2 norm of the filtered search statement by using the intermediate variable 1. Exemplarily, the text encoding module may calculate the first L2 norm of the filtered search statement according to 1 where β Represents the i-th element in T1.
[0236] Since the dimension of the image feature vector of the visual media is 768 dimensions, and in order to implement the calculation between the image feature vector and the text feature vector, while the intermediate variable 1 of the filtered search statement is 512 dimensions, therefore, the text encoding module needs to map the intermediate variable 1 of the filtered search statement to 768 dimensions. Then, the text encoding module can use the above mapping matrix 1 and the first L2 norm of the filtered search statement to map the intermediate variable 1 to a 768-dimensional text feature vector, so as to obtain the text feature vector of the filtered search statement corresponding to the base model. Specifically, the text encoding module can use To calculate the text feature vector of the filtered search statement corresponding to the base model. Among them, T’ 1 Represents the text feature vector of the filtered search statement corresponding to the base model, Is the transpose matrix of M 1 , and M 1 Represents the above mapping matrix 1.
[0237] For the fine-tuning model branch: The text encoding module uses the mapping matrix 4 and the feature vector of the filtered search statement to obtain the intermediate variable 2 of 512 dimensions. Exemplarily, the text encoding module calculates the intermediate variable 2 according to T 2 =YN 2 . Among them, T 2 Represents the intermediate variable 2, N 2 Represents the mapping matrix 4, and the text encoding module uses the intermediate variable 2 to calculate the second L2 norm of the filtered search statement. Exemplarily, the text encoding module can calculate according to To calculate the second L2 norm of the filtered search statement. Among them, β 2 Represents the second L2 norm of the filtered search statement, Represents the i-th element in T 2 .
[0238] Since the dimension of the image feature vector of the visual media is 768 dimensions, and in order to implement the calculation between the image feature vector and the text feature vector, while the intermediate variable 2 of the filtered search statement is 512 dimensions, therefore, the text encoding module needs to map the intermediate variable 2 of the filtered search statement to 768 dimensions. Then, the text encoding module can use the above mapping matrix 2 and the first L2 norm of the filtered search statement to map the intermediate variable 2 to a 768-dimensional text feature vector, so as to obtain the text feature vector of the filtered search statement corresponding to the fine-tuning model. Specifically, the text encoding module can use To calculate the text feature vector of the filtered search statement corresponding to the base model. Among them, T’ 2 Represents the text feature vector of the filtered search statement corresponding to the fine-tuning model, Is M 2The transposed matrix of, M 1 Represents the above mapping matrix 2.
[0239] It should be noted that the text feature vector of the filtered search statement corresponding to the fine-tuning model and the text feature vector of the filtered search statement corresponding to the base model can be output by the text encoding module simultaneously.
[0240] In the embodiments of the present application, the base model and the fine-tuning model in the custom CLIP model reuse the text encoder, so that the custom CLIP model only needs to calculate the feature vector of the filtered search statement once, and then can use the feature vector of the filtered search statement to respectively determine the text feature vector of the filtered search statement corresponding to the base model and the text feature vector of the filtered search statement corresponding to the fine-tuning model, without the base model and the fine-tuning model calculating the feature vector of the filtered search statement separately, thus improving the calculation efficiency of the text feature vector of the filtered search statement.
[0241] In the embodiments of the present application, in order to ensure the generalization ability of the multimodal model and the ability to recognize specific images and texts, the present application sets the multimodal model to include a base model and a fine-tuning model. Since the base model and the fine-tuning model have the same model weights, that is, part of the network is the same, directly setting two models will cause waste of mobile phone memory. Therefore, the present application reuses the image encoder and the text encoder for the base model and the fine-tuning model, so that the base model and the fine-tuning model can respectively use the feature vectors output by the image encoder and the text encoder to determine their respective corresponding text feature vectors and image feature vectors. Generally speaking, the determination time of the image feature vector and the text feature vector is reduced by nearly half.
[0242] S415. The multimodal understanding module performs recall based on the text feature vector and the image feature vector of the visual media to obtain candidate visual media.
[0243] In some embodiments, the image feature vector of the above visual media can be sent by the search module to the multimodal understanding module, or can be obtained by the multimodal understanding module from the local of the mobile phone.
[0244] In the embodiments of the present application, for each visual media on the mobile phone (such as each visual media in the above k visual media), the multimodal understanding module can calculate the vector similarity (or simply referred to as similarity) between the image feature vector of the visual media and the text feature vector of the filtered search statement, that is, calculate the similarity between the visual media and the filtered search statement based on the image feature vector of the visual media and the text feature vector of the filtered search statement. Then, the multimodal understanding module determines the visual media that matches the filtered search statement from the k visual media according to the vector similarity, and uses the determined visual media as the candidate visual media.
[0245] Exemplarily, the multimodal understanding module may determine a visual medium with a vector similarity greater than or equal to threshold 1 as a visual medium that matches the filtered search statement. Optionally, in a case where the number of visual media with a vector similarity greater than or equal to threshold 1 is greater than quantity 1, the multimodal understanding module may sort the visual media with a vector similarity greater than or equal to threshold 1 in descending order of vector similarity. Subsequently, the multimodal understanding module may determine the top n visual media as visual media that match the filtered search statement. Here, n is a positive integer.
[0246] In some embodiments, the image feature vector of the above-mentioned visual medium may include the image feature vector of the visual medium corresponding to the base model and the image feature vector of the visual medium corresponding to the fine-tuning model. The text feature vector of the above-mentioned filtered search statement includes the text feature vector of the filtered search statement corresponding to the base model and the text feature vector of the filtered search statement corresponding to the fine-tuning model. In one case, as Figure 5D shown, the multimodal understanding module may calculate the vector similarity 1 between the text feature vector of the filtered search statement corresponding to the base model and the image feature vector of the visual medium corresponding to the base model (that is, calculate the similarity between the visual medium corresponding to the base model and the filtered search statement), and calculate the vector similarity 2 between the text feature vector of the filtered search statement corresponding to the fine-tuning model and the image feature vector of the visual medium corresponding to the fine-tuning model (that is, calculate the similarity between the visual medium corresponding to the fine-tuning model and the filtered search statement). Subsequently, the multimodal understanding module may determine a visual medium with vector similarity 1 or vector similarity 2 greater than or equal to threshold 1 as a visual medium that matches the filtered search statement.
[0247] In another case, the multimodal understanding module may use whether the filtered search statement hits whitelist 1 to determine whether the filtered search statement is within the support range of the fine-tuning model, that is, determine whether to use the fine-tuning model branch to determine candidate visual media. The process of the multimodal understanding module using whitelist 1 to determine candidate visual media will be described below in conjunction with Figure 6 , to introduce the process of the multimodal understanding module using whitelist 1 to determine candidate visual media.
[0248] S501. The multimodal understanding module determines whether the filtered search statement belongs to whitelist 1.
[0249] In the embodiments of the present application, when the multimodal understanding module determines that the filtered search statement does not belong to the whitelist, it indicates that the filtered search statement is within the support range of the base model. The multimodal understanding module may use the base model branch to determine candidate visual media, and the multimodal understanding module may execute S502.
[0250] When the filtered search statement belongs to the whitelist, it indicates that the search statement is within the support scope of the fine-tuning model. The multimodal understanding module can enable the fine-tuning model branch to determine candidate visual media, and the multimodal understanding module can execute S504.
[0251] In some embodiments, the multimodal understanding module determines whether each semantic entity (i.e., visual semantic entity) in the filtered search statement belongs to Whitelist 1. Considering that the search statement input by the user is generally a phrase, the multimodal understanding module can determine that the filtered search statement belongs to the whitelist when each semantic entity of the filtered search statement belongs to Whitelist 1. When there is a semantic entity that does not belong to Whitelist 1, it is determined that the filtered search statement does not belong to the whitelist. For example, the filtered search statement is "a boy holding a basket", and the semantic entities include "basket" and "boy". The multimodal understanding module can respectively determine whether the basket and the boy belong to Whitelist 1. When both the basket and the boy belong to Whitelist 1, the multimodal understanding module can determine that the filtered search statement belongs to the whitelist. When either the basket or the boy does not belong to Whitelist 1, the multimodal understanding module can determine that the filtered search statement does not belong to the whitelist.
[0252] S502. For each visual media, the multimodal understanding module calculates the vector similarity 1 between the image feature vector of the visual media corresponding to the base model and the text feature vector of the filtered search statement corresponding to the base model.
[0253] In the embodiments of the present application, as Figure 7 shown, when the filtered search statement does not belong to Whitelist 1, the multimodal understanding module can call the text feature vector of the filtered search statement output by the base model branch of the text encoding module and the image feature vector of the visual media corresponding to the base model, and calculate the score corresponding to the visual media corresponding to the base model, that is, calculate the vector similarity 1 between the visual media corresponding to the base model and the filtered search statement.
[0254] Specifically, the multimodal understanding module can pass to calculate the score corresponding to the visual media corresponding to the base model. Among them, S 1 represents the score corresponding to the visual media corresponding to the base model, α1 represents the first L2 norm of the above-mentioned image feature vector of the visual media. X represents the above-mentioned image feature vector 1, and T’ 1 represents the text feature vector of the filtered search statement corresponding to the above-mentioned base model.
[0255] S503. The multimodal understanding module uses the visual media with a vector similarity 1 greater than the threshold 1 as the candidate visual media.
[0256] S504. For each visual medium, the multimodal understanding module calculates the vector similarity 2 between the image feature vector of the visual medium corresponding to the fine-tuning model and the text feature vector of the filtered search statement corresponding to the fine-tuning model.
[0257] In the embodiments of the present application, as described above Figure 7 As shown, when the filtered search statement belongs to the whitelist 1, the multimodal understanding module can call the text feature vector of the filtered search statement output by the fine-tuning model branch of the text encoding module and the image feature vector of the visual medium corresponding to the fine-tuning model, and calculate the score corresponding to the visual medium corresponding to the fine-tuning model, that is, calculate the vector similarity 2 between the visual medium corresponding to the fine-tuning model and the filtered search statement.
[0258] Specifically, the multimodal understanding module can pass through Calculate the score corresponding to the visual medium corresponding to the base model. Among them, S 2 Represents the score corresponding to the visual medium corresponding to the fine-tuning model, and α 2 Represents the second L2 norm of the image feature vector of the above visual medium. X represents the above image feature vector 1, and T’ 2 Represents the text feature vector of the filtered search statement corresponding to the above fine-tuning model.
[0259] S505. The multimodal understanding module uses the visual media with a vector similarity 2 greater than the threshold 1 as candidate visual media.
[0260] In some embodiments, the operations performed by the above multimodal understanding module can be performed by a multimodal model. In addition, the above S501-S505 can be performed by the output module of the multimodal understanding module, that is, the model in the multimodal model.
[0261] In the embodiments of the present application, the multimodal understanding module can determine the image feature vector of the visual medium offline, and only need to determine the text feature vector of the search statement online, which improves the calculation time of the vector similarity between the image feature vector and the text feature vector, thereby effectively reducing the user's retrieval time. In addition, the multimodal understanding module can determine whether to use the fine-tuning model to determine the candidate visual media according to whether the filtered search statement belongs to the whitelist 1, so as to improve the accuracy of determining the candidate visual media, thereby improving the accuracy of the search results.
[0262] Among them, optionally, the above whitelist 1 can be extended, that is to say, the support range of the fine-tuning model can be extended. In other words, the scenarios supported by the fine-tuning model can be extended. It only needs to train the fine-tuning model with the training set corresponding to the extended scenario. It should be understood that training the fine-tuning model is actually training the fine-tuning model branch in the above fine-tuning module and text encoding module, without training the reused parts (such as the above image encoder and text encoder).
[0263] In some embodiments, after determining the candidate visual media, the multimodal understanding module may directly use the above candidate visual media as the search result. After that, the multimodal understanding module may return the search result to the search module. Then, the search module may send the search result to the gallery service module for the gallery service module to display the search result.
[0264] In addition, to improve the accuracy of the search result, after obtaining the above candidate visual media, the mobile phone may continue to perform a secondary confirmation on the candidate visual media to further screen visual media from the candidate visual media to obtain visual media that matches the filtered search statement, i.e., the search statement. Optionally, the mobile phone may use both the visual media with vector similarity 1 greater than a threshold 1 (or referred to as the first threshold) and the visual media with vector similarity 2 greater than the threshold 1 as candidate visual media. In other words, the mobile phone may use the visual media determined by the base model and the fine-tuned model respectively as candidate visual media, that is, perform two searches. The process of secondary confirmation of the candidate visual media is introduced below.
[0265] S416. The multimodal understanding module sends the information of the filtered search statement and the image information of the candidate visual media to the secondary confirmation module.
[0266] The information of the filtered search statement includes one or more of the filtered search statement, the tokenization result of the filtered search statement, label 1 included in the filtered search statement, and the text feature vector of the filtered search statement.
[0267] The image information of the candidate visual media includes one or more of the similarity between the candidate visual media and the filtered search statement, label 2 of the candidate visual media, and the image feature vector of the candidate visual media.
[0268] In some embodiments, the above image feature vector and text feature vector are determined based on a multimodal model including a base model and a fine-tuned model. Correspondingly, the text feature vector of the filtered search statement may include the text feature vector of the filtered search statement corresponding to the base model and the text feature vector of the filtered search statement corresponding to the fine-tuned model. The image feature vector of the candidate visual media may include the image feature vector of the candidate visual media corresponding to the base model and the image feature vector of the candidate visual media corresponding to the fine-tuned model.
[0269] The similarity between the above candidate visual media and the filtered search statement may include the similarity between the candidate visual media corresponding to the base model and the filtered search statement (or referred to as similarity 1), and the similarity between the candidate visual media corresponding to the fine-tuned model and the filtered search statement (or referred to as similarity 2).
[0270] Of course, the above-mentioned image feature vectors, text feature vectors, and similarity may also be determined by only one model (such as the above-mentioned base model or fine-tuning model), and the present application does not limit this.
[0271] In some embodiments, the multimodal understanding module may perform word segmentation extraction on the filtered search statement to obtain the word segmentation result of the filtered search statement. For example, if the filtered search statement is "The child is playing by the sea", the word segmentation result is: child, is, by the sea, and playing. Exemplarily, the multimodal understanding module may perform word segmentation extraction on the filtered search statement through the natural language understanding engine service (the natural language understanding, NLU). Optionally, the multimodal understanding module may perform word segmentation extraction on the filtered search statement through the natural language understanding module.
[0272] After obtaining the word segmentation result of the filtered search statement, the multimodal understanding module may perform label mapping on the word segmentation result to obtain the label included in the filtered search statement, that is, label 1 (or referred to as the second label). Specifically, for each word in the word segmentation result, the filtering search module determines whether the word exists in the preset label table. In the case where the word does not exist in the preset label table, the multimodal understanding module may determine that there is no corresponding label for the word, that is, the filtered search statement does not include the label corresponding to the word.
[0273] In the case where the word exists in the preset label table, the multimodal understanding module may use the label corresponding to the word in the preset label table as the label 1 included in the filtered search statement. For example, if the word segmentation includes "cat", and the label corresponding to "cat" in the preset label table is "cat", then the label included in the filtered search statement includes "cat". It should be noted that the words in the word segmentation result of the filtered search statement and the labels corresponding to the words may be the same or different.
[0274] In some embodiments, the above-mentioned label 2 (or referred to as the first label) of the candidate visual media is obtained from the label library. The label 2 of the visual media in the label library represents the classification label of the visual media, which may be determined by the mobile phone (such as the multimodal understanding module in the mobile phone) using an image classification model to identify the visual media. Exemplarily, the classification label of the visual media may also be determined offline by the mobile phone.
[0275] S417. The secondary confirmation module determines the Boolean value corresponding to each candidate visual media based on the information of the filtered search statement and the image information of the candidate visual media.
[0276] S418. For each candidate visual media, when the Boolean value corresponding to the candidate visual media is true, the secondary confirmation module determines that the candidate visual media is visual media 1.
[0277] In the embodiments of the present application, the multimodal understanding module inputs the information of the filtered search statement and the image information of the candidate visual media into the secondary confirmation module, so that the secondary confirmation module determines the Boolean value corresponding to each candidate visual media, and the Boolean value corresponding to the candidate visual media indicates whether the candidate visual media is a visual media matching the filtered search statement, thereby realizing the secondary confirmation of the candidate visual media.
[0278] In some embodiments, the above-mentioned secondary confirmation module may include a binary classification module, a dynamic threshold module, a label confirmation module, and a whitelist threshold module. The binary classification module may use a binary classification model to determine whether there is a correlation between the filtered search statement and the candidate visual media, so as to obtain the Boolean value of the candidate visual media. The dynamic threshold module may use a dynamic threshold model to determine a threshold 2 that matches the length of the filtered search statement, so that the similarity between the candidate visual media and the filtered search statement can be compared with the threshold 2 to obtain the Boolean value of the candidate visual media. The label confirmation module may compare the label 2 of the candidate visual media with the label 1 included in the filtered search statement to obtain the Boolean value of the candidate visual media. The whitelist threshold module may determine a threshold 3 that matches the filtered search statement as a whole according to whether the filtered search statement hits a preset dictionary, so that the similarity between the candidate visual media and the filtered search statement can be compared with the threshold 3 to obtain the Boolean value of the candidate visual media. Among them, the detailed process of the secondary determination module determining the candidate visual media through the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module, that is, the implementation process of S417 above, can refer to the relevant descriptions below and will not be introduced here first.
[0279] S419. The secondary confirmation module returns visual media 1 to the search module.
[0280] S420. The search module obtains visual media 2 based on the non-visual semantic subject in the above search statement and the index of the visual media.
[0281] S421. The search module takes the intersection of visual media 1 and visual media 2 as the search result.
[0282] In the embodiments of the present application, the search module searches for the visual media (or referred to as visual media 2) corresponding to the filtered non-visual semantic subject in the search statement input by the user from the constructed index, that is, searches for the visual media whose attributes match the non-visual semantic subject, and uses it as visual media 2.
[0283] After that, the search module can determine the intersection of Visual Media 2 and Visual Media 1 to obtain the target visual media. The attributes of the target visual media match the non-visual semantic subject of the search statement, and the visual content of the target search media corresponds to the visual semantic subject of the search statement. For example, if the search statement is "the sky photographed this year", then Visual Media 1 includes sky content, and the acquisition time of Visual Media 2 is this year. Then, the target visual media includes sky content, and the acquisition time of the target visual media is this year.
[0284] S422. The search module sends the search result to the gallery service module.
[0285] S423. The gallery service module displays the search result.
[0286] In the embodiment of the present application, the gallery service module can display the search result to the user. For example, through Interface 305 above, Figure 3A in Interface 305 above, Figure 3A in Interface 308 above to display the search result. As shown in Interface 305 above, Figure 3A the search result includes 239 pictures, and Interface 305 only displays the thumbnails of 8 pictures in the search result; if the user wants to view these 239 pictures, they can click the "More" option in Interface 305, and in response to this operation, the mobile phone displays Interface 308 as shown in Figure 3A above.
[0287] In some embodiments, the gallery service module can sort the target visual media in the search result according to the acquisition time and display the target visual media. For example, sort the target visual media in ascending order of acquisition time, so that the earlier the acquisition time, the higher the display order.
[0288] In other embodiments, the gallery service module can sort the target visual media in the search result according to the similarity between the target visual media and the filtered search statement. For example, sort the target visual media in descending order of similarity, so that the target visual media with higher similarity has a higher display order, that is, the similarity between the target visual media displayed earlier and the filtered search statement is greater than or equal to the similarity between the target visual media displayed later and the filtered search statement. Here, the similarity between the target visual media and the filtered search statement can be Similarity 1 or Similarity 2, or a similarity calculated based on Similarity 1 and Similarity 2.
[0289] Optionally, when the filtered search statement belongs to the above-mentioned whitelist 1, the similarity between the target visual media and the filtered search statement can be similarity 2, that is, the similarity between the visual media corresponding to the fine-tuning model and the filtered search statement. When the filtered search statement does not belong to the above-mentioned whitelist 1, the similarity between the target visual media and the filtered search statement can be similarity 1, that is, the similarity between the visual media corresponding to the base model and the filtered search statement.
[0290] Optionally, the similarity calculated based on similarity 1 and similarity 2 can be the average of similarity 1 and similarity 2. Alternatively, calculate the weighted sum of similarity 1 and similarity 2, and this application does not limit it.
[0291] Next, a possible implementation process of the above S417 will be continued, as Figure 8 shown, this process may include S417a - S417g.
[0292] S417a. The secondary confirmation module determines whether the filtered search statement exists in the preset fine-tuning model vocabulary.
[0293] In the embodiments of this application, the secondary confirmation module can determine whether the filtered search statement exists in the preset fine-tuning model vocabulary (or called the first preset vocabulary) to determine whether the filtered search statement corresponds to the branch of the base model or the branch of the fine-tuning model, that is, to determine whether to perform secondary confirmation through the module corresponding to the base model or the module corresponding to the fine-tuning model.
[0294] When the filtered search statement does not exist in the above-mentioned preset fine-tuning vocabulary, it indicates that the filtered search statement corresponds to the branch of the base model, that is, it indicates that the candidate visual media is secondarily confirmed through the module corresponding to the base model to select a visual media that better matches the filtered search statement from the candidate visual media, and the secondary confirmation module can execute S417b.
[0295] When the filtered search statement exists in the above-mentioned preset fine-tuning vocabulary, it indicates that the filtered search statement corresponds to the branch of the fine-tuning model, that is, it indicates that the candidate visual media is secondarily confirmed through the module corresponding to the fine-tuning model to select a visual media that better matches the filtered search statement from the candidate visual media, and the secondary confirmation module can execute S417f.
[0296] S417b. The secondary confirmation module inputs the text feature vector of the filtered search statement and the image feature vectors of each candidate visual media into the binary classification module, and obtains the Boolean value 1 corresponding to each candidate visual media output by the binary classification module.
[0297] In the embodiments of this application, assume that the number of candidate visual media is m. As Figure 9AAs shown, the secondary confirmation module inputs the image feature vectors of each candidate visual medium among the m candidate visual media and the text feature vector of the filtered search statement into the binary classification model in the binary classification module. For each candidate visual medium among the m candidate visual media, the binary classification model determines the matching degree between the candidate visual medium and the search statement based on the image feature vector of the candidate visual medium and the text feature vector of the filtered search statement. This matching degree represents the degree of correlation between the candidate visual medium and the filtered search statement. The greater the matching degree, the more relevant the candidate visual medium is to the filtered search statement. Then, the binary classification model compares the matching degree between the candidate visual medium and the filtered search statement with the classification threshold (or called the preset classification threshold) to determine whether the candidate visual medium matches the search statement, obtaining the Boolean value (bool) 1 corresponding to the candidate visual medium, and thus outputting the Boolean values 1 corresponding to the m candidate visual media. When the matching degree is greater than the classification threshold, it indicates that the candidate visual medium matches the filtered search statement, and the Boolean value 1 corresponding to the candidate visual medium is true. When the matching degree is less than or equal to the classification threshold, it indicates that the candidate visual medium does not match the filtered search statement, and the Boolean value 1 corresponding to the candidate visual medium is false. Optionally, the above Boolean value 1 can also be used as the first matching value.
[0298] Among them, the binary classification model is a model pre-trained using a training sample set, which can determine whether these two match based on the image feature vector and the text feature vector. This binary classification model can be trained by the above mobile phone or other devices. Below, taking the binary classification model trained by a mobile phone as an example, the training process of the binary classification model will be introduced.
[0299] In some embodiments, when there is a filtered search statement in the above preset fine-tuning model vocabulary, it indicates that the filtered search statement corresponds to the branch of the fine-tuning model. Since the branch corresponding to the fine-tuning model does not include the binary classification module, the secondary confirmation module does not need to execute the above S417b.
[0300] When there is no segmentation of the filtered search statement in the above preset fine-tuning model vocabulary, it indicates that the filtered search statement corresponds to the branch of the base model. The text feature vector of the filtered search statement in the above S417b may include the text feature vector of the filtered search statement corresponding to the base model, and the image feature vector of the candidate visual medium in the above S417b may include the image feature vector of the candidate visual medium corresponding to the base model.
[0301] It can be understood that if the input parameters of the secondary confirmation module only include an image feature vector of a candidate visual medium and a text feature vector of a filtered search statement, and do not include the image feature vector of the candidate visual medium corresponding to the fine-tuning model and the image feature vector of the candidate visual medium corresponding to the base model, then the image feature vector of the candidate visual medium in S417b above is the input image feature vector of the candidate visual medium, and the text feature vector of the filtered search statement is the input text feature vector of the filtered search statement.
[0302] Exemplarily, the training sample set of the above binary classification model includes multiple training data, and each piece of training data in the multiple training data includes a sample image and its corresponding description text. The training sample set can include a positive sample training set (abbreviated as positive samples) and a negative sample training set (abbreviated as negative samples). The sample image in the training data included in the positive samples matches the description text corresponding to the sample image. For example, as Figure 9B shown in the image, the description text corresponding to this image is "a little boy holding a basket". Since Figure 9B the visual content expressed is that the little boy is holding a basket, which matches the corresponding description text, therefore, this image and its corresponding description text can be used as a piece of training data in the positive samples.
[0303] The sample image in the training data included in the negative samples does not match the description text corresponding to the sample image. For example, as Figure 9B shown in the image, the description text corresponding to this image is "flowers". Since Figure 9B the visual content expressed does not match the description text, therefore, this sample image and its corresponding description text can be used as a piece of training data in the negative samples.
[0304] It should be noted that since users generally use natural language when searching for visual media, the description text corresponding to the sample image also uses natural language, which conforms to the actual search situation.
[0305] To improve the training accuracy, it is necessary to ensure the quantity of the training data in the positive and negative samples. Considering that the efficiency of obtaining positive and negative samples manually is relatively low, the mobile phone can automatically generate positive and negative samples. Below, taking the above binary classification model as a text-image matching model as an example, the generation process of positive and negative samples will be continued. As Figure 10 shown, the process is as follows:
[0306] S1. The binary classification module obtains P sample data pairs. Among them, each sample data pair in the P sample data pairs includes a sample image and its corresponding description text.
[0307] S2. The binary classification module extracts sub-texts in the description text of each sample data.
[0308] Exemplarily, the above-mentioned sub-text may include descriptive phrases and / or nouns in the descriptive text. Exemplarily, the sub-text generally does not include adverbials, verbs, etc. in the descriptive text that do not correspond to actual visual content. The binary classification module can tokenize the descriptive text to determine the descriptive phrases and nouns in the descriptive text, and obtain the sub-text in the descriptive text. For example, if the descriptive text is "a little boy sitting in a basket", the sub-text in this descriptive text is "basket", "little boy", and "boy". Another example, the descriptive text is "a little boy sitting in a basket and sucking his thumb". After tokenizing this descriptive text, the obtained sub-text is "basket", "boy", "little boy", "little boy sucking his thumb", where "little boy sucking his thumb" can represent a descriptive phrase.
[0309] In some embodiments, the binary classification module can use the Han Language Processing Package (HanLP) to tokenize the descriptive text corresponding to the sample image.
[0310] S3. For each sample image, the binary classification module generates a text set corresponding to the sample image based on the descriptive text corresponding to the sample image and the sub-text in the descriptive text.
[0311] Among them, the text set corresponding to the sample image represents a description set of the visual content corresponding to the sample image. The text set corresponding to the sample image may include the descriptive text corresponding to the sample image and each sub-text in the descriptive text. For example, the descriptive text corresponding to the sample image is "a little boy holding a basket", and the sub-text in the descriptive text includes "basket", "little boy", and "boy". Correspondingly, the text set corresponding to this image is {"basket", "little boy", "boy", "a little boy holding a basket"}.
[0312] S4. The binary classification module randomly selects a text element from the text sets corresponding to each sample image to obtain the text element 1 corresponding to each sample image.
[0313] Among them, the text element in the text set can be a sub-text or a descriptive text.
[0314] In some embodiments, the binary classification module can randomly select text elements from the text sets corresponding to each sample image according to a preset ratio. The preset ratio includes the ratio of the text elements selected as descriptive texts and the ratio of the text elements selected as sub-texts. For example, the preset ratio is 8:2, the probability that the selected text element is a descriptive text is 80%, and the probability that the selected text element is a sub-text is 20%. Since generally the image encoder and the text encoder are trained using the image and its corresponding text as a whole, the ratio of the text elements selected as descriptive texts in the preset ratio will be greater than the ratio of the text elements selected as sub-texts to ensure the training effect of the model.
[0315] S5. For each sample image, the binary classification module respectively uses the P text elements 1 corresponding to the P sample images as the sample text elements corresponding to the sample image.
[0316] Among them, the number of sample text elements corresponding to one sample image among the P sample images is P. One sample image and one sample text element corresponding to it form a sample, so that multiple samples can be obtained. Since positive and negative samples are needed to train the model, after obtaining the samples, it is necessary to distinguish between positive and negative samples in the samples, so as to train the binary classification model using positive and negative samples. The process of distinguishing positive and negative samples in the samples will be introduced below.
[0317] In the embodiment of the present application, the binary classification module initially obtains P sample data pairs. These P sample data pairs are all positive samples. However, negative samples are also required for the training of the binary classification model. Therefore, the binary classification module can perform image-text matching on the P sample images and one text element (i.e., text element 1) in the text set corresponding to the P sample images. That is, for each sample image among the P sample images, the binary classification module can use all P text elements 1 as the sample text elements corresponding to the sample image. The sample image and each of the P text elements 1 form a sample, so that P*P samples can be obtained, realizing an increase in the number of samples, and thus realizing the rapid generation of samples. To implement the training of the binary classification model, the binary classification model needs to distinguish between positive and negative samples among the P*P samples. The process of distinguishing positive and negative samples will be continued to be introduced below.
[0318] S6. For each sample text element corresponding to the sample image, the binary classification module determines whether the sample text element belongs to the text set corresponding to the sample image.
[0319] In the embodiment of the present application, for each sample text element corresponding to the sample image, the binary classification module can determine whether the sample text element belongs to the text set corresponding to the sample image, so as to determine whether the sample text element corresponds to the visual content of the sample image, and thus determine whether the sample composed of the sample text element and the sample image is a positive sample.
[0320] In the case where the sample text element is in the text set corresponding to the sample image, it indicates that the sample text element corresponds to (i.e., matches) the visual content of the sample image. That is to say, the sample composed of the sample text element and the sample image is a positive sample. Therefore, the binary classification module can execute S7.
[0321] In the case where the sample text element is not in the text set corresponding to the sample image, it indicates that the sample text element does not correspond to (i.e., does not match) the visual content of the image. That is to say, the sample composed of the sample text element and the sample image is not a positive sample. Therefore, the binary classification module can execute S8.
[0322] S7. The binary classification module uses the sample image and the sample text element as positive samples.
[0323] S8. The binary classification module uses the image and the sample text element as negative samples.
[0324] In some embodiments, after obtaining the sample text elements corresponding to each sample image, the binary classification module may generate a corresponding sample matrix. Among them, the i-th row element in the sample matrix represents each sample text element corresponding to the i-th sample image, and the j-th column element in the sample matrix represents the text element 1 corresponding to the j-th sample image, that is, a text element randomly selected from the text set corresponding to the j-th sample image.
[0325] For example, the sample images include image a, image b, image c, and image d. Figure 11A The "1" in the sample matrix shown is the text element 1 corresponding to image a (i.e., the first sample image). And so on, "6" is the text element 1 corresponding to image d (i.e., the fourth sample image). The first row element 50 (i.e., 1, 2, 5, 6) is the sample text element corresponding to image a. The second column element 51 (i.e., 2, 2, 2, 2) is the text element 1 corresponding to image b.
[0326] Among them, a sample text element in the sample matrix and the sample image corresponding to the row where the sample text element is located form a sample. For example, as described above Figure 11A The "1" in the first row and first column shown forms a sample with image a.
[0327] In the embodiments of the present application, the binary classification model generates a corresponding sample matrix by mutually matching P sample images and P text elements 1, so that the binary classification module can distinguish whether the sample text elements in the sample matrix belong to positive samples or negative samples, realizing the rapid determination of positive and negative samples, improving the determination efficiency of positive and negative samples, and moreover, by generating a label matrix, the comprehensiveness of image-text matching can be ensured, avoiding the situation of missing sample images or text elements 1, such as avoiding the situation of not matching a certain sample image with a certain text element 1.
[0328] After obtaining the sample matrix, for each sample text element in the sample matrix, the binary classification module needs to determine whether the sample text element matches the sample image corresponding to the row where the sample text element is located. To improve the matching efficiency, the binary classification module can use the text intersection matrix 1 to match with the sample matrix to obtain positive and negative samples. The text intersection matrix 1 represents the intersection of the text sets corresponding to the sample images. Exemplarily, the determination process of the text intersection matrix 1 may include:
[0329] The binary classification module determines the overlapping elements in the text sets corresponding to any two sample images. Then, the binary classification module generates a corresponding text intersection matrix 1. Among them, the s-th element in the t-th row of the text intersection matrix 1 represents the overlapping elements in the text sets between the t-th sample image and the s-th sample image.
[0330] For example, the text set corresponding to image a (as Figure 11B shown) is {1, 2, 3}, the text set corresponding to image b is {2, 3, 4}, the text set corresponding to image c is {5, 6}, and the text set corresponding to image d is {6, 7}. It should be understood that the numbers here are actually corresponding text elements (such as the above sub-texts, description texts).
[0331] Then, the binary classification module can determine the overlapping elements (i.e., the same elements) in the text sets corresponding to any two sample images among image a, image b, image c, and image d. For example, the same elements in the text sets corresponding to image a and image b are 2, 3.
[0332] Then, the binary classification module generates a text intersection matrix 1 as Figure 11C shown. This text intersection matrix 1 includes the overlapping elements in the text sets corresponding to the sample image and each sample image (i.e., image a, image b, image c, and image d). Among them, Figure 11C the element in the first row and first column of the text intersection matrix 1 shown represents the overlapping elements in the text set between image a and image a (i.e., the text set corresponding to image a). The element in the first row and second column represents the overlapping elements in the text set between image a and image b. And so on, Figure 11C the element in the fourth row and fourth column in
[0333] In some embodiments, after the binary classification module obtains the text sets corresponding to each sample image among the P sample images, it can calculate the union of the text sets corresponding to each sample image. The union of this text set can include the text elements in all text sets. Then, the binary classification module can assign numbers (such as the above numbers) to each text element in the union of the text sets, so that different text elements in the text set correspond to different numbers, and the same text elements in different text sets correspond to the same number. By assigning numbers to text elements, the graphic-text matching efficiency can be improved, and thus the generation efficiency of positive and negative samples can be improved.
[0334] The determination process of the text intersection matrix 1 is introduced above. Next, the process of using the text intersection matrix 1 to match with the sample matrix to determine positive and negative samples will be continued.
[0335] The binary classification module intersects the above-mentioned text intersection matrix 1 and the sample matrix to obtain text intersection matrix 2.
[0336] After that, for each intersection element in text intersection matrix 2, the binary classification module determines whether the intersection element is empty.
[0337] In the case where the intersection element is not empty, it indicates that the visual content of the intersection element and the sample image corresponding to its row match. Therefore, the binary classification module can confirm that the intersection element and the sample image corresponding to its row are positive samples.
[0338] In the case where the intersection element is empty, it indicates that the visual content of the intersection element and the sample image corresponding to its row do not match. Therefore, the binary classification module can confirm that the element and the sample image corresponding to its row are negative samples, thus achieving the rapid determination of positive and negative samples. Among them, the element corresponding to the sample image can also be called the text corresponding to the sample image.
[0339] Optionally, the binary classification module can use 1 and 0 to distinguish whether the intersection elements in text intersection matrix 2 belong to positive samples or negative samples, that is, to distinguish whether the sample text elements in the sample matrix belong to positive samples or negative samples. In the case where the intersection element in text intersection matrix 2 is empty, the binary classification module can set the label corresponding to the intersection element to 0, that is, label the sample text element corresponding to the intersection element with 0.
[0340] In the case where the intersection element in text intersection matrix 2 is not empty, the binary classification module can set the label corresponding to the element to 1, that is, label the sample text element corresponding to the intersection element with 1, thereby obtaining the positive and negative sample label matrix and realizing the setting of labels for the sample text elements in the sample matrix. The position of the intersection element in text intersection matrix 2 is the same as the position of the sample text element corresponding to the intersection element in the sample matrix.
[0341] Based on this, the binary classification module can use the sample text element corresponding to the label 0 and the image corresponding to the row where the sample text element is located as a piece of training data in the negative samples, and use the sample text element corresponding to the label 1 and the image corresponding to the row where the sample text element is located as a piece of training data in the positive samples, thus achieving the batch determination of multiple positive and negative sample training data.
[0342] For example, as Figure 11D shown, the text intersection matrix 1 and the sample matrix are intersected to obtain text intersection matrix 2. After that, the binary classification module can determine whether the intersection elements in text intersection matrix 2 are empty, thereby determining the positive and negative sample label matrix corresponding to the sample matrix. As Figure 11DThe label in the first row and first column of the positive and negative sample label matrix is 1. The element in the sample matrix corresponding to this label is the sample text element "1" in the first row and first column. This "1" and image a form a positive sample. As Figure 11D The label in the second row and first column of the positive and negative sample label matrix is 0. The sample text element in the second row and first column of the sample matrix corresponding to this label is "1". This "1" and image b form a negative sample.
[0343] It should be noted that generally, the elements on the diagonal of the sample matrix and the sample images corresponding to their rows form positive samples.
[0344] In the embodiments of the present application, the binary classification module intersects the text intersection matrix 1 with the sample matrix to obtain the text intersection matrix 2, so as to quickly determine positive and negative samples by using whether the elements in the text intersection matrix 2 are empty, improve the determination efficiency of positive and negative samples, and ensure the comprehensiveness and integrity of the distinction between positive and negative samples, so that the model can be trained quickly.
[0345] The process of determining positive and negative samples is introduced above. Next, the process of training a binary classification model using positive and negative samples will be continued.
[0346] S9. The binary classification module uses positive and negative samples to train the image-text matching model to obtain the trained image-text matching model.
[0347] Exemplarily, the image-text matching model can be a multilayer perceptron (MLP) model. As Figure 11E Shown, the binary classification module uses a text encoder to process the sample text elements in positive and negative samples to obtain text feature vectors of the sample text elements. And, the binary classification module uses an image encoder to process the sample images in positive and negative samples to obtain image feature vectors of the sample images. Then, for each sample in positive and negative samples, the binary classification module can splice the text feature vector of the sample text element in the sample and the image feature vector of the sample image in the sample. Then, the binary classification module can input the spliced text feature vector and image feature vector into the MLP model to train the MLP model. During the training process, the binary classification module can use a loss function to test the difference between the predicted value output by the trained MLP model and the actual value. The predicted value represents whether the predicted sample image and text match. The actual value represents the actual matching situation between the sample image and text.
[0348] In some embodiments, the above loss function can be a focal loss function. Specifically, the focal loss function can be FL(p t )=-at 1(1 - p t ) γ log(p t ). Wherein, pt represents the probability of the matching between the predicted sample image and its corresponding text by the image - text matching model, that is, the probability that the predicted sample image and its corresponding text are positive samples. at1 is a factor for adjusting the weights of positive and negative samples, which can be set according to the number of positive and negative samples to adjust the imbalance of the number of positive and negative samples. For example, at1 is 0.1. γ is a adjustment factor used to reduce the loss contribution of easily distinguishable samples. It should be understood that the larger pt is, the smaller the value of the loss function is.
[0349] Optionally, the above - mentioned image encoder can be a stacked auto - encoder, and the text encoder can be a count vectorizer.
[0350] It should be noted that the above - introduced focal loss function is only an example of the loss function. The loss function can also be other types of loss functions. For example, the loss function is a loss function of the BCE class. In addition, the above - mentioned MLP model is only an example of the image - text matching model. The image - text matching model can also be other deep learning models, and the present application does not limit it.
[0351] In some embodiments, the above - mentioned positive - negative sample label matrix can be utilized when calculating the value of the loss function.
[0352] In some embodiments, the use of the image - text matching model to re - confirm the candidate visual media is only an example. The image - text matching model can also be directly used for searching visual media. For example, after the user inputs a search statement, the image - text matching model can directly use the text feature vector of the search statement (or filtered search statement) and the image feature vectors of each visual media on the mobile phone to determine whether the visual media matches the search statement, and obtain the corresponding boolean value.
[0353] The process of using the binary classification module to determine whether the candidate visual media matches the filtered search statement is introduced above. Next, in combination with S417c, the process of using the dynamic threshold module to determine whether the candidate visual media matches the filtered search statement will be introduced.
[0354] S417c: The binary classification module inputs the filtered search statement and the similarity between each candidate visual media and the filtered search statement into the dynamic threshold module, and obtains the boolean value 2 corresponding to each candidate visual media output by the dynamic threshold module.
[0355] In the embodiments of the present application, generally, the similarity threshold 1 corresponding to search statements of different lengths is different. The longer the length of the search statement, the higher the similarity degree between the visual content of the visual media and the search statement is required to be, and correspondingly, the similarity threshold 1 needs to be higher. Therefore, the secondary confirmation module can use the dynamic threshold module to determine a dynamic threshold that matches the length of the filtered search statement, that is, to determine the similarity threshold 1 (or referred to as the first similarity threshold) corresponding to the filtered search statement, so as to determine the Boolean value 2 corresponding to each candidate visual media by using the similarity threshold 1 corresponding to the filtered search statement. Optionally, the Boolean value 2 can also be used as the second matching value.
[0356] Exemplarily, as Figure 12 shown, the process by which the dynamic threshold module determines the Boolean value 2 corresponding to each candidate visual media may include: First, the dynamic threshold module can determine the length of the filtered search statement through the filtered search statement. After that, based on the length of the filtered search statement, the dynamic threshold model combines t = parameter 1 * L + parameter 2 to determine the similarity threshold 1 corresponding to the filtered search statement. Among them, the above parameters 1 and 2 are pre-set parameters. For example, parameter 1 is 0.05 and parameter 2 is 0.33. Optionally, the dynamic threshold module can also determine the length of the filtered search statement through the word segmentation result of the filtered search statement.
[0357] After that, for each candidate visual media among the m candidate visual media, the dynamic threshold module can compare the size between the similarity between the candidate visual media and the filtered search statement and the similarity threshold 1 corresponding to the filtered search statement. In the case where the similarity between the candidate visual media and the filtered search statement is less than the similarity threshold 1, it indicates that the similarity degree between the visual content of the candidate visual media and the filtered search statement is relatively low, and the dynamic threshold module can determine that the Boolean value 2 corresponding to the candidate visual media is false.
[0358] In the case where the similarity between the candidate visual media and the filtered search statement is greater than or equal to the similarity threshold 1, it indicates that the similarity degree between the visual content of the candidate visual media and the filtered search statement is relatively high, and the dynamic threshold module can determine that the Boolean value 2 corresponding to the candidate visual media is true.
[0359] Among them, optionally, as described above Figure 12 shown, the value range of t can be greater than or equal to parameter 3, that is, min(t, parameter 3), and the value range of t can be less than or equal to parameter 4. That is, max(t, parameter 4). Among them, parameters 3 and 4 are preset.
[0360] In some embodiments, since the dynamic threshold module corresponding to S417c belongs to the branch corresponding to the base model, the similarity between the candidate visual media and the filtered search statement in S417c includes the similarity between the candidate visual media corresponding to the base model and the filtered search statement (i.e., similarity 1). Correspondingly, the dynamic threshold module can determine whether similarity 1 between the candidate visual media and the filtered search statement is less than similarity threshold 1. When similarity 1 is greater than or equal to similarity threshold 1, the dynamic threshold module can determine that the Boolean value 2 corresponding to the candidate visual media is true. In the case where similarity 1 is less than similarity threshold 1, the dynamic threshold module can determine that the Boolean value 2 corresponding to the candidate visual media is false.
[0361] It can be understood that if the input parameters of the secondary confirmation module only include one similarity between the candidate visual media and the filtered search statement, rather than the similarity between the candidate visual media corresponding to the fine-tuning model and the filtered search statement and the similarity between the candidate visual media corresponding to the base model and the filtered search statement, then the image feature vector of the candidate visual media in S417c is the similarity between the input candidate visual media and the filtered search statement.
[0362] In addition, the similarity between the candidate visual media and the filtered search statement can also be determined by other multi-modal models, and the present application does not limit it.
[0363] The process of determining the Boolean value corresponding to the candidate visual media using the dynamic threshold module is introduced above. Next, the process of determining whether the candidate visual media matches the filtered search statement using the label confirmation module will be continued with reference to S417d.
[0364] S417d: The secondary confirmation module inputs the word segmentation result of the above filtered search statement, label 1 included in the filtered search statement, and label 2 of each candidate visual media into the label confirmation module, and obtains the Boolean value 3 corresponding to each candidate visual media output by the label confirmation module.
[0365] In the embodiments of the present application, the secondary confirmation module can use the label confirmation module to determine whether label 2 of the candidate visual media exists in the word segmentation of the filtered search statement or label 1 included in the filtered search statement, so as to determine whether the candidate visual media matches the filtered search statement, and thus determine the Boolean value 3 corresponding to the candidate visual media. Optionally, the Boolean value 3 can also be used as the third matching value.
[0366] Exemplarily, as Figure 13 shown, the process of the label confirmation model determining the Boolean value 3 corresponding to the candidate visual media can include:
[0367] First, for each of the m candidate visual media, the label confirmation module can determine whether the label 2 of the candidate visual media includes the word segments in the word segmentation result of the filtered search statement, and determine whether the label 2 of the candidate visual media includes the label 1 included in the filtered search statement, that is, determine whether the candidate visual media includes the visual content corresponding to the filtered search statement.
[0368] In the case where the label 2 of the candidate visual media includes at least one word segment of the filtered search statement, or the label 2 of the candidate visual media includes at least one label 1 included in the filtered search statement, it indicates that the candidate visual media hits the filtered search statement. The label confirmation module can determine that the candidate visual media may be the visual media required by the user. Therefore, the label confirmation module can determine the Boolean value 3 corresponding to the candidate visual media as true. For example, the label 1 of the filtered search statement includes "child" and "flower", and the word segmentation result of the filtered search statement includes the word segments "child", "holding", and "flower". In the case where the label 2 of the candidate visual media includes "child", "holding" or "flower", or the label 2 of the candidate visual media includes "child" or "flower", it is determined that the Boolean value 3 corresponding to the candidate visual media is true.
[0369] In the case where the label 2 of the candidate visual media does not include all the word segments of the filtered search statement, and the label 2 of the candidate visual media does not include all the label 1 included in the filtered search statement, it indicates that the candidate visual media does not hit the filtered search statement. The candidate visual media may not be the visual media required by the user. Therefore, the label confirmation module can determine the Boolean value 3 corresponding to the candidate visual media as false, so as to obtain the Boolean value 3 of the m candidate visual media.
[0370] Optionally, in order to improve the accuracy of the search results, the label confirmation module can determine that the Boolean value 3 corresponding to the candidate visual media is true when the label 2 of the candidate visual media includes all the word segments of the filtered search statement, or includes all the label 1 included in the filtered search statement.
[0371] In the case where the label 2 of the candidate visual media does not include at least one word segment of the filtered search statement and does not include at least one label 1 included in the filtered search statement, it is determined that the Boolean value 3 corresponding to the candidate visual media is false.
[0372] The process of the secondary confirmation module using the label confirmation module to determine the Boolean value 3 corresponding to each candidate visual media is introduced above. Next, the process of the secondary determination module using the whitelist threshold module to determine whether the candidate visual media matches the filtered search statement will be continued.
[0373] The secondary confirmation module inputs the similarity between the filtered search statement and the candidate visual media into the whitelist threshold module, and obtains the Boolean value 4 corresponding to each candidate visual media output by the whitelist threshold module.
[0374] In the embodiments of the present application, the secondary confirmation module can determine whether the filtered search statement hits the dictionary through the whitelist threshold module, so as to determine whether there is a similarity threshold 2 corresponding to the filtered search statement, and accurately determine the similarity threshold. Exemplarily, the whitelist threshold module can determine whether the dictionary contains the filtered search statement. If it exists, the whitelist threshold module becomes effective, and the whitelist threshold module can use the similarity threshold corresponding to the filtered search statement in the dictionary as the similarity threshold 2 corresponding to the filtered search statement, so as to screen the candidate visual media matching the filtered search statement by using the similarity threshold 2. If it does not exist, the whitelist threshold module does not become effective, that is, there is no need to use the whitelist threshold module to determine the Boolean value 4 corresponding to each candidate visual media. Among them, the dictionary includes a key and the value corresponding to the key. The key represents a preset search text, and the value is the similarity threshold corresponding to the preset search text. The key in the dictionary can be the search text frequently input by the user, that is, the search statement that has been pre-tested. Optionally, the Boolean value 4 can also be used as the fourth matching value.
[0375] In some embodiments, there is a corresponding dictionary for the branches corresponding to the fine-tuning model and the base model. As Figure 14A shown, when the filtered search statement is the branch corresponding to the base model, the whitelist threshold module can determine whether the dictionary corresponding to the base model (or called the second preset word list) includes the filtered search statement, that is, determine whether there is a key in the dictionary corresponding to the base model that is the same as the filtered search statement.
[0376] When the dictionary corresponding to the base model includes the filtered search statement, it indicates that the filtered search statement hits the dictionary corresponding to the base model, that is, there is a key in the dictionary corresponding to the base model that is the same as the filtered search statement. Then, the whitelist threshold module can use the value corresponding to the filtered search statement in the dictionary corresponding to the base model as the similarity threshold 2.
[0377] After that, for each of the m candidate visual media, the whitelist threshold module may compare the similarity between the candidate visual media corresponding to the base model and the filtering search threshold with the similarity threshold 2. In the case where the similarity between the candidate visual media corresponding to the base model and the filtering search threshold is less than the similarity threshold 2, the whitelist threshold module may determine that the Boolean value 4 corresponding to the candidate visual media is false. In the case where the similarity is greater than or equal to the similarity threshold 2, indicating that the visual content of the candidate visual media is highly similar to the filtering search statement, the whitelist threshold module may determine that the Boolean value corresponding to the candidate visual media is true.
[0378] Correspondingly, in the case where the filtering search statement is the branch corresponding to the base model, the above S418 may be that for each candidate visual media, in the case where the Boolean value corresponding to the candidate visual media is true, the secondary confirmation module (such as the fusion module in the secondary confirmation module) may use the candidate visual media as the visual media 1. Among them, the Boolean value corresponding to the candidate visual media includes the above Boolean value 1, Boolean value 2, Boolean value 3, and Boolean value 4. That is to say, the secondary confirmation module may take the union of the Boolean value 1, Boolean value 2, Boolean value 3, and Boolean value 4 corresponding to the candidate visual media to obtain the Boolean value corresponding to the candidate visual media.
[0379] In the case where the Boolean values corresponding to the candidate visual media are all false, indicating that the Boolean value 1, Boolean value 2, Boolean value 3, and Boolean value 4 corresponding to the candidate visual media are all false, the secondary confirmation module (such as the fusion module in the secondary confirmation module) may not use the candidate visual media as the visual media 1.
[0380] It should be noted that since both the above whitelist threshold module and the above dynamic threshold module are used to confirm the similarity threshold corresponding to the filtering search statement, and the similarity threshold determined by the whitelist threshold module is more accurate. Therefore, when the above whitelist threshold module becomes effective, that is, when using the similarity threshold 2 corresponding to the filtering search statement to determine the Boolean value corresponding to the candidate visual media, the dynamic threshold module may not become effective, and the secondary confirmation module may not use the similarity threshold 1 corresponding to the filtering search statement to determine the Boolean value corresponding to the candidate visual media, avoiding unnecessary determination of the similarity threshold and thus avoiding unnecessary screening of candidate visual media.
[0381] In some embodiments, the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module in the above secondary confirmation module may determine the Boolean value corresponding to the candidate visual media in parallel or serially. Whether it is determined in parallel or serially, when a module in the secondary confirmation module determines that the Boolean value corresponding to the candidate visual media is true, other modules do not need to continue to determine the Boolean value corresponding to the candidate visual media, that is, there is no need to input the relevant information of the candidate visual media into other models, avoiding unnecessary determination of Boolean values and ensuring the transmission cost.
[0382] In addition, the above secondary confirmation module including the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module is only an example. The secondary confirmation module may include one or more of the binary classification module, the dynamic threshold module, the label confirmation model, and the whitelist threshold module to improve the search efficiency of visual media. For example, if the secondary confirmation module includes one of the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module, such as the secondary confirmation module includes the binary classification module, correspondingly, the Boolean value corresponding to the above candidate visual media may include Boolean value 1. Another example is that the secondary confirmation module includes two of the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module, such as the secondary confirmation module includes the binary classification module and the dynamic threshold module. Correspondingly, the Boolean value corresponding to the above candidate visual media may include Boolean value 1 and Boolean value 2. Another example is that the secondary confirmation module includes three of the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module, such as the secondary confirmation module includes the binary classification module, the dynamic threshold module, and the label confirmation module. Correspondingly, the Boolean value corresponding to the above candidate visual media may include Boolean value 1, Boolean value 2, and Boolean value 3.
[0383] Moreover, the content included in the image information of the above candidate visual media and the information of the filtering search statement, that is, the input parameters of the secondary confirmation module are only an example, and can be adaptively set according to the modules included in the secondary confirmation module. For example, if the secondary confirmation module includes a binary classification model, the image information of the above candidate visual media may include the image feature vector of the candidate visual media, and the information of the filtering search statement may include the text feature vector of the filtering search statement.
[0384] The above introduced the process in which the secondary confirmation module can sequentially use the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module to determine the Boolean value corresponding to the candidate visual media when the filtering search statement corresponds to the branch of the base model. Next, the process in which the secondary confirmation module can sequentially use the label confirmation module and the whitelist threshold module to determine the Boolean value corresponding to the candidate visual media when the filtering search statement corresponds to the branch of the fine-tuning model will be continued.
[0385] S417f. The secondary confirmation module inputs the word segmentation result of the above filtering search statement, label 1 included in the filtering search statement, and label 2 of each candidate visual media into the label confirmation module, and obtains the Boolean value 5 corresponding to each candidate visual media output by the label confirmation module.
[0386] Among them, the implementation process of S417f can refer to the implementation process of the above S417d, and will not be elaborated here. Optionally, the Boolean value 5 can also be used as the fifth matching value.
[0387] S417g. The secondary confirmation module inputs the filtering search statement, the similarity between the candidate visual media and the filtering search statement into the whitelist threshold module, and obtains the Boolean value 6 corresponding to each candidate visual media output by the whitelist threshold module.
[0388] Among them, the implementation process of S417g can refer to the implementation process of the above S417e. As Figure 14B shown, when the filtering search statement is the branch corresponding to the fine-tuning model, the whitelist threshold module can determine whether the dictionary corresponding to the fine-tuning model includes the filtering search statement, that is, determine whether there is the same key as the filtering search statement in the dictionary (or called the third preset word list) corresponding to the fine-tuning model, so as to determine whether the whitelist threshold module takes effect.
[0389] After the whitelist threshold module takes effect, the secondary confirmation module can compare the similarity between the candidate visual media corresponding to the fine-tuning model and the filtering search statement with the value corresponding to the filtering search statement in the dictionary of the fine-tuning model to determine the Boolean value corresponding to the candidate visual media. Optionally, the Boolean value 6 can also be used as the sixth matching value.
[0390] Correspondingly, when the filtering search statement is the branch corresponding to the fine-tuning model, the above S418 can be that for each candidate visual media, when the Boolean value corresponding to the candidate visual media is true, the secondary confirmation module (such as the fusion module in the secondary confirmation module) can use the candidate visual media as the visual media 1. Among them, the Boolean value corresponding to the candidate visual media includes the above Boolean value 5 and the above Boolean value 6, that is to say, the secondary confirmation module can take the union of the Boolean value 5 and the Boolean value 6 corresponding to the candidate visual media to obtain the Boolean value corresponding to the candidate visual media.
[0391] In the embodiments of the present application, after obtaining the candidate visual media, the mobile phone uses the secondary confirmation module to continue to determine the Boolean value corresponding to the candidate visual media, so as to determine whether the candidate visual media matches the filtered search statement, so that the candidate visual media that matches the filtered search statement can be selected from the candidate visual media, and the corresponding search results can be obtained, ensuring the accuracy of the search results, and thus ensuring the user satisfaction.
[0392] It should be noted that the operations performed by the above modules or models are only examples. The operations performed by the above modules can also be performed by other modules in the mobile phone, and the present application does not limit it. In addition, the operations actually performed by the above modules or models are performed by the mobile phone.
[0393] In some embodiments, when the mobile phone determines the candidate visual media or performs secondary confirmation on the candidate visual media, it may also not filter the non-visual semantic entities in the search statement first, but directly use the text feature vector of the search statement to determine the candidate visual media or perform secondary confirmation on the candidate media.
[0394] In some embodiments, the search and storage of the above visual media are carried out under user authorization, including but not limited to notifying and reminding the user to read the relevant user agreement (notification) before the user uses this function, and signing the agreement (authorization) including authorizing relevant user information.
[0395] In some embodiments, the present application provides a computer storage medium, including computer instructions, when the computer instructions run on an electronic device, enabling the electronic device to execute the method as described above.
[0396] In some embodiments, the present application provides a computer program product, when the computer program product runs on an electronic device, enabling the electronic device to execute the method as described above.
[0397] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part.
[0398] It should be understood that the "embodiments" mentioned throughout the specification mean that the specific features, structures or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, the embodiments throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0399] It should also be understood that in the present application, "when", "if" and "in case" all mean that under certain objective circumstances, the UE or the base station will perform corresponding processing, which does not limit the time, and does not require the UE or the base station to have a judgment action when implemented, nor does it mean that there are other limitations.
[0400] Those of ordinary skill in the art can understand that the various numerical numbers such as the first, second, etc. involved in the present application are only for the convenience of description and do not limit the scope of the embodiments of the present application, nor do they represent the order of precedence.
[0401] In the present application, the elements represented by the singular are intended to mean "one or more", rather than "one and only one", unless otherwise specified. In the present application, unless otherwise specified, "at least one" is intended to mean "one or more", and "a plurality" is intended to mean "two or more".
[0402] The term "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. These three situations, where A can be singular or plural, and B can be singular or plural.
[0403] The term "at least one of..." or "at least one kind of..." in this article means all or any combination of the items listed. For example, "at least one of A, B and C" can represent: A exists alone, B exists alone, C exists alone, A and B exist simultaneously, B and C exist simultaneously, and A, B and C exist simultaneously. These six situations, where A can be singular or plural, B can be singular or plural, and C can be singular or plural.
[0404] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0405] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0406] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0407] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0408] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0409] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0410] The same or similar parts among the various embodiments in this application can be referred to each other. In the various embodiments in this application, as well as in each implementation manner / implementation method / realization method in each embodiment, if there is no special explanation and logical conflict, the terms and / or descriptions among different embodiments, as well as among each implementation manner / implementation method / realization method in each embodiment, are consistent and can be mutually referred to. The technical features in different embodiments, as well as in each implementation manner / implementation method / realization method in each embodiment, can be combined to form new embodiments, implementation manners, implementation methods, or realization methods according to their internal logical relationships. The above-described implementation manners of this application do not constitute a limitation on the protection scope of this application.
[0411] As described above, the above are only the specific implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights. In short, the above is only a preferred embodiment of the technical solution of this application and is not used to limit the protection scope of this application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.
Claims
1. A visual media search method applicable to an electronic device, characterized in that, the method includes: displaying a first interface; the first interface includes a search box; receiving a search statement input in the search box; determining candidate visual media; wherein, the candidate visual media represents visual media with a similarity greater than a first threshold to the search statement; determining a matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the search statement; determining target visual media according to the matching value corresponding to the candidate visual media; displaying search results, the search results corresponding to the target visual media.
2. The method according to claim 1, characterized in that, the method further includes: in the case where the first preset word list does not include the search statement, the matching value corresponding to the candidate visual media is determined according to a first target matching value corresponding to the candidate visual media, the first target matching value includes at least one of a first matching value, a second matching value, a third matching value and a fourth matching value; the first matching value is determined according to the matching degree between the candidate visual media and the search statement; the second matching value is determined by comparing the similarity between the candidate visual media and the search statement and a first similarity threshold corresponding to the search statement; the third matching value is determined based on whether the visual content in the candidate visual media includes the visual content corresponding to the search statement; the fourth matching value is determined by comparing the similarity between the candidate visual media and the search statement and a second similarity threshold corresponding to the search statement; in the case where the first preset word list includes the search statement, the matching value corresponding to the candidate visual media is determined according to a second target matching value corresponding to the candidate visual media, the second target matching value includes a fifth matching value and / or a sixth matching value; the fifth matching value is determined based on whether the visual content in the candidate visual media includes the visual content corresponding to the search statement, and the sixth matching value is determined by comparing the similarity between the candidate visual media and the search statement and a third similarity threshold corresponding to the search statement.
3. The method according to claim 2, characterized in that, the image information of the candidate visual media includes the image feature vector of the candidate visual media, and the information of the search statement includes the text feature vector of the search statement; the determining the matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the search statement includes: inputting the image feature vectors of the candidate visual media and the text feature vector of the search statement into a binary classification model to obtain a first matching value corresponding to each candidate visual media; the binary classification model is used to determine the matching degree between the candidate visual media and the search statement based on the image feature vector of the candidate visual media and the text feature vector of the search statement, and determine the first matching value corresponding to the candidate visual media according to the matching degree and a preset classification threshold.
4. The method according to claim 3, wherein, the image feature vector of the candidate visual media includes the image feature vector of the candidate visual media corresponding to the base model, and the text feature vector of the search statement includes the text feature vector of the search statement corresponding to the base model.
5. The method according to any one of claims 2 to 4, wherein, the image information of the candidate visual media includes the similarity between the candidate visual media and the search statement; determining the matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the search statement includes: determining a first similarity threshold corresponding to the search statement based on the length of the search statement; when the similarity between the candidate visual media and the search statement is greater than or equal to the first similarity threshold, determining that the second matching value corresponding to the candidate visual media is true; when the similarity between the candidate visual media and the search statement is less than the first similarity threshold, determining that the second matching value corresponding to the candidate visual media is false.
6. The method according to claim 5, wherein, the similarity between the candidate visual media and the search statement includes the similarity between the candidate visual media corresponding to the base model and the search statement.
7. The method according to any one of claims 2 to 6, wherein, the image information of the candidate visual media includes a first label of the candidate visual media; the first label represents the classification to which the visual content included in the candidate visual media belongs; the information of the search statement includes the word segmentation of the search statement and a second label; the second label of the search statement is obtained by mapping the word segmentation of the search statement; determining the matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the search statement includes: when the first label of the candidate visual media includes the word segmentation of the search statement or includes the second label of the search statement, determining that the third matching value corresponding to the candidate visual media is true; when the first label of the candidate visual media does not include the word segmentation of the search statement and does not include the second label of the search statement, determining that the third matching value corresponding to the candidate visual media is false.
8. The method according to any one of claims 2 to 7, wherein, the image information of the candidate visual media includes the similarity between the candidate visual media and the search statement, and the information of the search statement includes the search statement; determining the matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the search statement includes: When the second preset vocabulary includes the search statement, if the similarity between the candidate visual media and the search statement is greater than or equal to the second similarity threshold, it is determined that the fourth matching value corresponding to the candidate visual media is true; the second similarity threshold refers to the similarity threshold corresponding to the search statement in the second preset vocabulary. If the similarity between the candidate visual media and the search statement is less than the second similarity threshold, it is determined that the fourth matching value corresponding to the candidate visual media is false.
9. The method according to claim 2, wherein, the image information of the candidate visual media includes the first label of the candidate visual media; the information of the search statement includes the word segmentation and the second label of the search statement. The determining of the matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the search statement includes: when the first label of the candidate visual media includes the word segmentation of the search statement or includes the second label of the search statement, it is determined that the fifth matching value corresponding to the candidate visual media is true; when the first label of the candidate visual media does not include the word segmentation of the search statement and does not include the second label of the search statement, it is determined that the fifth matching value corresponding to the candidate visual media is false.
10. The method according to claim 2 or 9, wherein, the image information of the candidate visual media includes the similarity between the candidate visual media and the search statement, and the information of the search statement includes the search statement. The determining of the matching value corresponding to the candidate visual media based on the image information of the candidate visual media and the information of the search statement includes: when the third preset vocabulary includes the search statement, if the similarity between the candidate visual media and the search statement is greater than or equal to the third similarity threshold, it is determined that the sixth matching value corresponding to the candidate visual media is true; the third similarity threshold refers to the similarity threshold corresponding to the search statement in the third preset vocabulary; if the similarity between the candidate visual media and the search statement is less than the third similarity threshold, it is determined that the sixth matching value corresponding to the candidate visual media is false.
11. The method according to claim 10, wherein, the similarity between the candidate visual media and the search statement refers to the similarity between the candidate visual media corresponding to the fine-tuning model and the search statement; the text feature vector of the search statement refers to the text feature vector of the search statement corresponding to the fine-tuning model.
12. The method according to any one of claims 1 to 11, wherein, the determining of the target visual media according to the matching value corresponding to the candidate visual media includes: when the matching value corresponding to the candidate visual media is true, it is determined that the candidate visual media is the target visual media.
13. The method according to any one of claims 2 to 12, wherein, When the first target matching value corresponding to the candidate visual media includes true, the matching value corresponding to the candidate visual media is true; When the first target matching value corresponding to the candidate visual media does not include true, the matching value corresponding to the candidate visual media is false.
14. The method according to any one of claims 1 to 13, wherein, the image information of the candidate visual media includes at least one of the similarity between the candidate visual media and the search statement, the image feature vector of the candidate visual media, and the first label of the candidate visual media; the information of the search statement includes at least one of the search statement, the text feature vector of the search statement, the second label included in the search statement, and the word segmentation of the search statement.
15. The method according to any one of claims 1 to 14, wherein, the determining of the candidate visual media includes: for each visual media on the electronic device, determining the similarity between the visual media and the search statement based on the image feature vector of the visual media and the text feature vector of the search statement; when the similarity is greater than the first threshold, taking the visual media as the candidate visual media.
16. The method according to any one of claims 1 to 15, wherein, the matching value corresponding to the candidate visual media is determined based on the image information of the candidate visual media and the information of the filtered search statement: the filtered search statement is obtained by filtering non-visual semantic entities in the search statement, and the non-visual semantic entities do not correspond to visual content.
17. An electronic device, wherein, the electronic device includes a display screen, a memory, and one or more processors; the display screen, the memory, and the processors are coupled; the display screen is used to display images generated by the processors, the memory is used to store computer program code, and the computer program code includes computer instructions; when the processors execute the computer instructions, the electronic device executes the method according to any one of claims 1 to 16.
18. A computer storage medium, wherein, including computer instructions, when the computer instructions run on an electronic device, the electronic device executes the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Visual media personalized search method and device
CN113641857A
Video searching method and device, electronic equipment and storage medium
CN113901330A
Text adjusted visual search
US20220138247A1