Visual media searching method and electronic equipment

By combining the basic model and the fine-tuning model, using the training set of preset search scenarios to train the fine-tuning model, the problem of inaccurate retrieval of complex visual media search statements in the prior art is solved, and high-accurate visual media search is achieved.

CN120067393AActive Publication Date: 2025-05-30HONOR DEVICE CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202311581946.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-30
Estimated Expiration
2043-11-23

AI Technical Summary

Technical Problem

The prior art is difficult to accurately search complex visual media search statements, resulting in inaccurate search results.

Method used

The method of combining basic model and fine-tuning model is adopted to train fine-tuning models through the training set of preset search scenarios to improve the accuracy of similarity calculation between visual media and search statements.

Benefits of technology

Accurate retrieval of complex search statements is achieved, and the accuracy and efficiency of visual media search is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067393A_ABST
    Figure CN120067393A_ABST
Patent Text Reader

Abstract

The invention provides a visual media searching method and electronic equipment, and relates to the technical field of image processing. After receiving a search statement input by a user, the electronic equipment determines a text feature vector of the search statement. And then, the electronic equipment can judge whether the search statement hits the first preset white list or not. Under the condition that the first preset white list is not hit, it is indicated that the search statement corresponds to the base model branch, and the electronic device can calculate the similarity between the text feature vector of the search statement and the image feature vector of the visual media corresponding to the base model by means of the text feature vector of the search statement and the image feature vector. Under the condition that the first preset white list is hit, it is indicated that the search statement corresponds to the fine tuning model branch, and the electronic device can calculate the similarity between the text feature vector of the search statement and the image feature vector of the visual media corresponding to the fine tuning model; therefore, the electronic equipment can determine the search result by utilizing the similarity between the visual media and the search statement, and accurate retrieval of the visual media is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to a visual media search method and an electronic device. Background Art

[0002] With the development of electronic devices (such as mobile phones), the shooting function of mobile phones has also developed rapidly. More and more users use mobile phones to take photos, videos, etc., and store the taken photos and videos in the gallery of the mobile phone. In addition, users can also store the pictures and screenshots downloaded by the mobile phone in the gallery of the mobile phone.

[0003] When a user wants to search for visual media (such as photos, videos, etc.), the user can enter a search statement on the mobile phone. For example, for photos taken on September 1st, the mobile phone responds to the search statement input by the user and retrieves the photos taken on September 1st to obtain corresponding search results. However, the mobile phone's ability to understand search statements is limited. When the search statement is relatively complex, the mobile phone may not be able to accurately obtain the corresponding search results. Summary of the Invention

[0004] In view of this, this application provides a visual media search method and an electronic device for improving the accuracy of search results.

[0005] In a first aspect, this application provides a visual media search method applied to an electronic device. The electronic device can display a first interface, and the first interface includes a search box. The electronic device can receive a search statement input by the user in the search box. Thereafter, the electronic device can use, as first visual media, visual media with a first similarity to the search statement greater than a first threshold, or visual media with a second similarity to the search statement greater than a second threshold. Herein, the first similarity between the visual media and the search statement is the similarity determined by the base model between the visual media and the search statement, that is, the first similarity is the similarity between the visual media corresponding to the base model and the search statement. The first similarity between the visual media and the search statement is the similarity determined by the fine-tuning model between the visual media and the search statement, that is, the first similarity is the similarity between the visual media corresponding to the fine-tuning model and the search statement.

[0006] The above-mentioned fine-tuning model is obtained by training the base model based on a training set corresponding to a preset search scenario. The fine-tuning model has a corresponding support range. When the search statement falls within this support range, the fine-tuning model can accurately retrieve visual media matching the search statement. The training set corresponding to the preset search scenario may include training images corresponding to the preset search scenario and description texts corresponding to the training images.

[0007] Thereafter, the electronic device can display search results corresponding to the search statement, and the search results include the first visual media.

[0008] In this application, the electronic device can use the base model or the fine-tuned model to determine the similarity between the visual media and the search statement, so that the electronic device can use the visual media with a higher similarity as the first visual media matching the search statement, realizing the accurate search of visual media. Whether the search statement is complex or not, the accurate search of the first visual media can be realized to meet the user's search needs. Moreover, since the fine-tuned model is trained on the basis of the base model using the training set corresponding to a specific search scenario, the fine-tuned model can be used to calculate the similarity between the search statement and the visual media in a specific search scenario, ensuring the accuracy of the similarity calculation.

[0009] In a possible design, the electronic device can determine whether to use the base model or the fine-tuned model to determine the similarity according to whether the search statement belongs to a preset search scenario, that is, whether it belongs to the first preset whitelist. When the search statement does not belong to the first preset whitelist (or called whitelist 1), it indicates that the search statement does not belong to the preset search scenario. The electronic device can use the visual media with a first similarity greater than the first threshold (or called threshold 1) to the search statement as the first visual media. When the search statement belongs to the first preset whitelist, it indicates that the search statement belongs to the preset search scenario. The electronic device can use the visual media with a second similarity greater than the second threshold to the search statement as the first visual media. Based on this, the electronic device can determine whether to use the base model or the fine-tuned model to determine the similarity between the visual media and the search statement according to the search statement, that is, according to the demand, ensuring the accuracy of the similarity determination.

[0010] Optionally, the above second threshold may be the same value as the above first threshold or may be a different value, and this application does not limit it.

[0011] In a possible design, the process by which the electronic device determines whether the above search statement belongs to the above first preset whitelist may include:

[0012] The electronic device can determine whether the first preset whitelist includes the whole search statement. When the first preset whitelist includes the search statement, the electronic device can determine that the search statement belongs to the first preset whitelist. Otherwise, the electronic device can determine that the search statement does not belong to the first preset whitelist.

[0013] Or,

[0014] The electronic device can determine whether each semantic entity in the search statement belongs to the first preset whitelist. When each semantic entity in the search statement belongs to the first preset whitelist, the electronic device can determine that the search statement belongs to the first preset whitelist, that is, the fine-tuned model branch corresponding to the search statement.

[0015] In the case where the semantic entity in the search statement does not belong to the first preset whitelist, the electronic device can determine that the search statement does not belong to the first preset whitelist, that is, the base model branch corresponding to the search statement. Based on this, the electronic device can accurately judge the model branch, so as to accurately calculate the similarity.

[0016] In a possible design, the first similarity between the above visual media and the above search statement is obtained based on the first image feature vector of the visual media determined by the base model and the first text feature vector of the search statement determined by the base model; specifically, the first similarity between the visual media and the search statement refers to the vector similarity between the image feature vector of the visual media corresponding to the base model and the text feature vector of the search statement corresponding to the base model.

[0017] The second similarity between the above visual media and the above search statement is obtained based on the second image feature vector of the visual media determined by the fine-tuning model and the second text feature vector of the search statement determined by the fine-tuning model; specifically, the second similarity between the visual media and the search statement refers to the vector similarity between the image feature vector of the visual media corresponding to the fine-tuning model and the text feature vector of the search statement corresponding to the fine-tuning model. Based on this, the first similarity determined by the base model is obtained based on the image feature vector determined by the base model and the text feature vector determined by the base model, ensuring the accuracy of the first similarity determined by the base model. The second similarity determined by the fine-tuning model is obtained based on the image feature vector determined by the fine-tuning model and the text feature vector determined by the fine-tuning model, ensuring the accuracy of the second similarity determined by the fine-tuning model.

[0018] In a possible design, the first image feature vector of the above visual media, and the second image feature vector of the visual media can be determined in advance when the electronic device is in an offline state, such as the screen-off charging state. The electronic device only needs to determine the text feature vector of the search statement online, which can improve the search efficiency of the visual media.

[0019] In a possible design, the above base model and fine-tuning model can reuse the image encoder. The first image feature vector of the above visual media can include the first feature vector of the visual media output by the image encoder (or referred to as image feature vector 1) and the first L2 norm of the visual media, and the first L2 norm of the visual media is determined based on the first feature vector of the visual media and the first mapping matrix (or referred to as mapping matrix 1) in the base model.

[0020] The second image feature vector includes the first feature vector and the second L2 norm of the visual media, and the second L2 norm of the visual media is determined based on the first feature vector of the visual media and the second mapping matrix (or referred to as mapping matrix 2) in the fine-tuning model.

[0021] In this application, by reusing the image encoder in the base model and the fine-tuning model, the electronic device can determine the first feature vector only once, instead of the base model determining the first feature vector once and the fine-tuning model determining the first feature vector once. Instead, the base model and the fine-tuning model can directly utilize the first feature vector output by the image encoder respectively, improving the determination efficiency of the first image feature vector and the second image feature vector of the visual media. And the resource occupancy can be reduced.

[0022] In a possible design, the above-mentioned base model and fine-tuning model reuse the text encoder. The first text feature vector of the visual media is determined based on the second feature vector of the search statement (or referred to as the feature vector for filtering the search statement) output by the text encoder and the third mapping matrix (or referred to as mapping matrix 3) in the base model. The second text feature vector of the visual media is determined based on the second feature vector and the fourth mapping matrix (or referred to as mapping matrix 4) in the fine-tuning model.

[0023] In this application, by reusing the text encoder in the base model and the fine-tuning model, the electronic device can determine the second feature vector only once, instead of the base model determining the second feature vector once and the fine-tuning model determining the second feature vector once. Instead, the base model and the fine-tuning model can directly utilize the second feature vector output by the text encoder respectively, improving the determination efficiency of the first text feature vector and the second text feature vector of the search statement. And the resource occupancy can be reduced.

[0024] In a possible design, the electronic device can save the first feature vector, the first L2 norm, and the second L2 norm of the visual media in the electronic device to implement the saving of the first image feature vector and the second image feature vector.

[0025] Among them, the dimension of the above-mentioned first feature vector is 768 dimensions. Therefore, the electronic device can save the first image feature vector and the second image feature vector with 770 dimensions, instead of saving 512 * 2 dimensions, reducing the resource occupancy.

[0026] In a possible design, the first similarity between the above-mentioned visual media and the search statement can be adopted to determine; S 1 represents the first similarity between the visual media and the search statement, and α 1 represents the first L2 norm of the visual media, and T′ 1The first text feature vector representing the search statement, and X represents the first feature vector of the visual media.

[0027] The second similarity between the above visual media and the search statement can be adopted Determined; S 2 Represents the second similarity between the visual media and the search statement, α 2 Represents the second L2 norm of the visual media, T′ 2 The second text feature vector representing the search statement, and X represents the first feature vector of the visual media.

[0028] Optionally, the above first visual media can be used as a candidate visual media. The electronic device continues to perform a secondary confirmation on the candidate visual media.

[0029] In a possible design, after obtaining the search statement, the electronic device can filter the non-visual semantic entities in the search statement to obtain a filtered search statement. Then, the electronic device can use the visual media with a first similarity greater than a first threshold to the filtered search statement as the third visual media, or use the visual media with a second similarity greater than a second threshold to the filtered search statement as the third visual media. And, the electronic device can determine a second visual media (or referred to as visual media 2) that matches the non-visual semantic entity in the search statement.

[0030] After that, the electronic device can use the intersection between the third visual media and the second visual media as the first visual media to accurately determine the search result.

[0031] Among them, the above first mapping matrix and third mapping matrix are obtained by training a base model with training set 1. When training the first mapping matrix, the images in training set 1 can be used for training. When training the second mapping matrix, the text corresponding to the images in training set 1 can be used for training. Similarly, the above second mapping matrix and fourth mapping matrix are obtained by training a fine-tuning model with a training set under a preset search scenario. When training the second mapping matrix, the images in this training set can be used for training. When training the fourth mapping matrix, the text corresponding to the images in this training set can be used for training.

[0032] In a possible design, in response to an operation of opening the gallery application, display the first interface;

[0033] Or, in response to an operation of opening the negative first screen triggered on the main screen of the electronic device, display the first interface;

[0034] Or, in response to a pull-down search operation triggered on the main screen of the electronic device, display the first interface.

[0035] In a second aspect, the present application provides an electronic device, which includes a display screen, a memory, and one or more processors; the display screen, the memory, and the processors are coupled; the display screen is configured to display an image generated by the processor, the memory is configured to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the electronic device is caused to execute the method as described above.

[0036] In a third aspect, the present application provides a computer storage medium, including computer instructions, which when running on an electronic device, cause the electronic device to execute the method as described above.

[0037] In a fourth aspect, the present application provides a computer program product, which when running on an electronic device, causes the electronic device to execute the method as described above.

[0038] It can be understood that for the beneficial effects that can be achieved by the electronic device described in the second aspect, the computer storage medium described in the third aspect, and the computer program product described in the fourth aspect provided above, reference can be made to the beneficial effects in the first aspect and any of its possible design manners, and details are not elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1A FIG. 1 is a schematic diagram of an interface for visual media search provided by an embodiment of the present application;

[0040] Figure 1B FIG. 2 is a schematic diagram of an interface for visual media search provided by an embodiment of the present application;

[0041] Figure 1C FIG. 3 is a schematic diagram of an interface for visual media search provided by an embodiment of the present application;

[0042] Figure 1D FIG. 4 is a schematic diagram of an interface for visual media search provided by an embodiment of the present application; Figure Four ;

[0043] Figure 2A FIG. 5 is a block diagram of the structure of an electronic device provided by an embodiment of the present application;

[0044] Figure 2B FIG. 6 is a software structure diagram of an electronic device provided by an embodiment of the present application;

[0045] Figure 3A FIG. 7 is a schematic diagram of an interface for visual media search provided by an embodiment of the present application;

[0046] Figure 3B FIG. 8 is a schematic diagram of an interface for visual media search provided by an embodiment of the present application;Figure Six ;

[0047] Figure 4 Schematic diagram 1 of a visual media search method provided by an embodiment of the present application;

[0048] Figure 5A Schematic diagram 1 of a process for determining an image feature vector provided by an embodiment of the present application;

[0049] Figure 5B Schematic diagram 2 of a process for determining an image feature vector provided by an embodiment of the present application;

[0050] Figure 5C Schematic diagram of a process for determining a text feature vector provided by an embodiment of the present application;

[0051] Figure 5D Schematic diagram of a process for determining a similarity provided by an embodiment of the present application;

[0052] Figure 6 Schematic diagram 2 of a visual media search method provided by an embodiment of the present application;

[0053] Figure 7 Schematic diagram 3 of a visual media search method provided by an embodiment of the present application;

[0054] Figure 8 Schematic diagram of a visual media search method provided by an embodiment of the present application Figure Four ;

[0055] Figure 9A Schematic diagram 1 of a secondary confirmation provided by an embodiment of the present application;

[0056] Figure 9B Schematic diagram 1 of a visual media provided by an embodiment of the present application;

[0057] Figure 10 Schematic diagram 5 of a visual media search method provided by an embodiment of the present application;

[0058] Figure 11A Schematic diagram 1 of a matrix provided by an embodiment of the present application;

[0059] Figure 11B Schematic diagram 2 of a visual media provided by an embodiment of the present application;

[0060] Figure 11C Schematic diagram 2 of a matrix provided by an embodiment of the present application;

[0061] Figure 11D Schematic diagram of a positive and negative sample discrimination process provided by an embodiment of the present application;

[0062] Figure 11E Schematic diagram of the training process of a binary classification model provided by an embodiment of the present application;

[0063] Figure 12 Schematic diagram two of a secondary confirmation provided by an embodiment of the present application;

[0064] Figure 13 Schematic diagram three of a secondary confirmation provided by an embodiment of the present application;

[0065] Figure 14A Schematic illustration of a secondary confirmation provided by an embodiment of the present application Figure Four ;

[0066] Figure 14B Schematic diagram five of a secondary confirmation provided by an embodiment of the present application. Detailed implementation manners

[0067] Next, the technical solutions in the embodiments of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application. Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this embodiment, unless otherwise specified, the meaning of "a plurality" is two or more.

[0068] To understand the embodiments of the present application more clearly, the following first explains the vocabulary involved in the present application.

[0069] Visual media: refers to pictures or videos.

[0070] Semantic subject: Named entity recognition technology (NER) can identify statements and recognize entities with specific meanings in the text, such as personal names and place names. In this solution, the entities with specific meanings identified are called semantic subjects.

[0071] Visual content related and visual content unrelated: Visual content refers to the objects presented by visual media and their interrelationships, etc. Briefly, visual content can be understood as the content included in visual media. In the context of image search in the present application, the data that can be obtained after the natural picture understanding of visual media files by the model is called "visual content related". In this solution, the data that is related to visual media files and can be obtained without the picture understanding ability of the model is called "visual content unrelated". For example, when an electronic device collects visual media files, it can obtain and save the shooting location, shooting time, name, file attributes, etc.

[0072] For example, in the phrase "photos taken in City 1 this year", "this year" (the time of shooting), "City 1" (the location of shooting), and "photos" (the file attribute) are all data that can be obtained and saved by the electronic device when collecting visual media files. Therefore, "this year", "City 1", and "photos" have nothing to do with visual semantics. In the phrase "the sky in City 1 taken this year", "the sky" can only be obtained through the image understanding ability of the model to understand the picture. Therefore, "the sky" is related to visual semantics.

[0073] Text semantic vector: It can be obtained by feeding the text into a text encoder, which is a vector that can represent the semantic features of the entire sentence. The text encoder can use the clip model or other models, such as the Transformer commonly used in natural language processing (NLP). This solution is not limited here. Among them, in this application, the text semantic vector can also be called the text feature vector.

[0074] Visual semantic vector: It can be obtained by feeding visual media (such as images) into an image encoder. The image encoder can use the clip model or other models, such as the CNN model or the VIT model. This solution is not limited here. Among them, in this application, the visual semantic vector can also be called the image feature vector.

[0075] Vector similarity: It is used to describe the similarity between two vectors (for example, between the text semantic vector and the visual semantic vector). In the embodiments of this application, the similarity between the text semantic vector of the search statement and the visual semantic vector of the visual media can be compared to determine the visual media that matches the search statement. Generally, the vector similarity can be calculated through the cosine similarity calculation formula. Of course, it can also be calculated through other methods.

[0076] An electronic device (such as a mobile phone) can manage the user's pictures, videos and other visual media through the gallery application. Taking the example of taking a photo with a mobile phone, after the mobile phone takes a photo, the gallery application determines the attribute tags such as the shooting location, shooting time, and photo name corresponding to the photo and can use the attribute tags as the index of the picture. After the gallery application establishes an index for the visual media, it can provide the corresponding search service to the user. Specifically, the user can search for pictures or videos on the mobile phone by entering keywords in the gallery application. Exemplarily, the user can enter keywords such as "sky", "cat", "time point 1" in the search box provided by the gallery application, and the gallery application matches the keywords entered by the user with the indexes of the pictures, videos and other visual media in the gallery application to obtain the search results.

[0077] Optionally, the above-mentioned attribute tags may further include attributes such as the face identity number (industrialdesign, ID) of the person in the photo, the person's name, and the relationship between the person and the mobile phone user. Among them, the face ID of the person in the photo can be automatically generated by the gallery application, and the person's name in the photo and the relationship between the person and the mobile phone user can be manually input by the user. In practical applications, the same person corresponds to the same face ID, the same name, and the same relationship with the mobile phone user. Therefore, in order to simplify the user operation, the user only needs to input the person's name and the relationship with himself once for the same person. Subsequently, the gallery application automatically configures the person's face ID, name, and the relationship between the person and the mobile phone user for the pictures containing the person's face through face recognition technology. In addition, the above-mentioned attribute information such as the photo name can also be automatically generated by the gallery application or manually named by the user.

[0078] The following exemplarily describes the interface involved in the search process of the gallery application with reference to the accompanying drawings:

[0079] As Figure 1A shown in (a) of Figure 1A , the mobile phone can display the main interface 101, which can also be called the desktop. The main interface 101 may include the icon 102 of the gallery application. The mobile phone receives the operation of the user clicking the icon 102. In response to this operation, the mobile phone can start the gallery application and display the interface 103 as shown in (b) of

[0080] As Figure 1A shown in (b) of

[0081] As Figure 1A shown in (b) of Figure 1AThe interface 105 shown in (c) therein can be called a search interface. Among them, the interface 105 can display classification information of photos to the user. For example, in the interface 105, the mobile phone classifies the photos of the local machine according to time, portraits, and things, etc. For example, in the dimension of time, the mobile phone classifies the photos of the local machine according to three time periods: "this month", "last month", and "this year". Among them, the "this month" album includes the photos or videos taken by the mobile phone this month, the "last month" album includes the photos or videos taken by the mobile phone last month, and the "this year" album includes the photos or videos taken by the mobile phone this year. In the dimension of portraits, the mobile phone classifies the photos of the local machine according to different people, such as the four different people in the interface 105. In the dimension of things, the mobile phone classifies and displays the photos of the local machine according to "scenery", "animals", "documents", and "buildings". It should be noted that the above classification dimensions can also be others, and no specific restrictions are made here. In the interface 105, the user can see this classification information without entering keywords.

[0082] Optionally, the interface 105 can also include options for search history 107 and "clear" 108. The search history includes the keywords that the user has entered, such as "flowers", "coffee", "cats", etc. The mobile phone can receive the operation of the user clicking "clear" 108. In response to this operation, the mobile phone can clear the search history. After the mobile phone clears the search history, the keywords that the user has entered are no longer displayed on the search interface 105. For example, in response to the operation of the user clicking "clear" 108, as Figure 1B shown, the options for search history 107 and "clear" 108 are no longer displayed on the search interface 105, and the content displayed below moves up.

[0083] In response to the operation of the user entering the keyword "sky" in the interface 105, the mobile phone displays as Figure 1CThe interface 109 shown in (a) therein. Among them, the mobile phone can search for data related to the keyword "sky" on the local device. Specifically, the mobile phone can associate with the keyword "sky" to obtain associated words such as "sky" and photos containing the word "sky". Then, search according to each associated word to obtain the search results of each associated word. For example: 100 photos related to "sky" and 32 photos related to photos containing the word "sky". Among them, the 100 photos related to "sky" can be recalled because the classification labels of each of these 100 photos match "sky" or "sky" and its associated words; the 32 photos related to photos containing the word "sky" can be recalled because through optical character recognition (OCR) technology, it is recognized that these 32 photos contain characters such as "sky". The union of the search results of each of these multiple associated words can be used as the search result of the keyword "sky".

[0084] The interface 109 also displays some search results of the keyword "sky" and a "More" option 110 corresponding to the search result of the keyword "sky". The mobile phone receives the click operation of the user on the "More" option 110 and displays as Figure 1C the interface 111 shown in (b) therein. Among them, the interface 111 is used to display photos and videos in the search results of the keyword "sky". Optionally, the photos and videos can be classified and displayed according to time. In addition, the interface 111 also includes a return key 112 and a title 113. In response to the user's operation on the return key 112, the mobile phone can redisplay the interface 109. The title 113 may include the keyword "sky".

[0085] That is to say, in the gallery application, when the user enters a simple search statement in the search box, such as a simple keyword, for example: sky, location 1, time 1, etc., corresponding search results can be obtained. However, because the mobile phone's understanding and association ability of the search statement is limited, if the user enters a more complex search statement in the search box, if the keywords in the search statement cannot match the attribute labels of the pictures or the text in the pictures, no photos may be found. That is to say, the existing gallery applications do not support the search function based on complex search statements. As Figure 1D shown, when the user enters a more complex search statement "warming oneself by the fire and boiling tea" in the search box of the interface 114, the mobile phone cannot understand the associated words of "warming oneself by the fire and boiling tea", and since the photos do not have labels that can match "warming oneself by the fire and boiling tea" or its associated words, the search result shows "no pictures".

[0086] In some embodiments, an electronic device may utilize an encoder to calculate the similarity between visual media on the electronic device and a search statement, so that the electronic device can use the visual media with a similarity higher than a threshold as search results. After that, the electronic device displays the search results to present the visual media required by the user. For example, the encoder may include a text encoder and an image encoder. The electronic device may input each visual media on the electronic device into the image encoder to obtain the image feature vectors corresponding to each visual media. And the electronic device may input the search statement into the text encoder to obtain the text feature adjacent to the search statement. After that, for each visual media, the electronic device may calculate the similarity between the visual media and the search statement based on the image feature vector corresponding to the visual media and the text feature vector corresponding to the search statement.

[0087] For another example, the above encoder may be a single encoder. The electronic device may input each visual media and the search statement on the electronic device into the encoder respectively, so that the encoder can obtain the similarity between each visual media and the search statement according to the image feature vectors of each visual media and the text feature vectors of the search statement. However, the electronic device may not be able to accurately determine the image feature vectors of the visual media only by using the encoder, resulting in a low calculation accuracy of the similarity between the visual media and the search statement, thus causing the search results displayed by the electronic device may not be the visual media required by the user and unable to meet the user's search needs. For example, when the search statement includes landmark buildings, the encoder may not be able to accurately determine the image feature vectors of the visual media including landmark buildings, resulting in a low accuracy of the calculated similarity between the visual media and the search statement, and further causing the electronic device unable to accurately search for the visual media required by the user. Simply put, the electronic device cannot accurately identify the visual media including landmark buildings, resulting in a low accuracy of the determined search results. For another example, when the search statement includes rare animals, the encoder may not be able to accurately identify the visual media including rare animals, resulting in a low accuracy of the determined search results.

[0088] Therefore, in order to improve the accuracy of visual media search to meet the search needs of users, the present application provides a visual media search method. After receiving a search statement input by a user, an electronic device determines a text feature vector of the search statement online. Then, the electronic device can determine whether the search statement hits a first preset whitelist to determine whether the search statement corresponds to a base model branch or a fine-tuning model branch. In the case of not hitting the first preset whitelist, it indicates that the search statement corresponds to the base model branch. For each visual media stored in the electronic device, the electronic device can use the text feature vector of the search statement and the image feature vector of the visual media corresponding to the base model to calculate the similarity between the visual media and the search statement. In the case of hitting the first preset whitelist, it indicates that the search statement corresponds to the fine-tuning model branch. For each visual media stored in the electronic device, the electronic device can use the text feature vector of the search statement and the image feature vector of the visual media corresponding to the fine-tuning model to calculate the similarity between the two. Thus, the electronic device can use the similarity between the visual media and the search statement to determine the search result, realizing the accurate retrieval of visual media and meeting the search needs of users. Moreover, the above-mentioned image feature vector of the visual media can be determined in advance by the electronic device, that is, determined offline, without the need to be determined only after the user inputs the search statement, improving the search efficiency of visual media.

[0089] Exemplarily, the above-mentioned electronic device can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., which can store visual media.

[0090] Exemplarily, Figure 2A FIG. shows a schematic structural diagram of an electronic device 200. The electronic device 100 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone jack 270D, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.

[0091] The sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.

[0092] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0093] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0094] The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0095] A memory may also be provided in the processor 210 for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory may save the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0096] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0097] It can be understood that the interface connection relationships among the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0098] The electronic device 200 realizes the display function through a GPU (Graphics Processing Unit), a display screen 294, and an application processor, etc. The GPU is connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change display information.

[0099] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 294, where N is a positive integer greater than 1.

[0100] The electronic device 200 can implement the shooting function through the ISP, the camera 293, the video codec, the GPU, the display screen 294, and the application processor, etc.

[0101] The camera 293 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In some embodiments, the electronic device 100 may include one or N cameras 293, where N is a positive integer greater than 1.

[0102] The video codec is used to compress or decompress digital videos. The electronic device 200 can support one or more video codecs. In this way, the electronic device 200 can play or record videos in multiple encoding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0103] The NPU is a neural-network (NN) computing processor. By drawing on the structure of the biological neural network, such as the transmission pattern between human brain neurons, it can quickly process the input information and can also continuously learn on its own. Through the NPU, applications such as the intelligent cognition of the electronic device 200 can be realized, such as: image recognition, face recognition, voice recognition, text understanding, etc.

[0104] The external memory interface 220 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 200. The external memory card communicates with the processor 210 through the external memory interface 220 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.

[0105] The internal memory 221 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 221 can include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.). The data storage area can store the data created during the use of the electronic device 200 (such as audio data, phone book, etc.). In addition, the internal memory 221 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. The processor 210 executes various functional applications and data processing of the electronic device 200 by running the instructions stored in the internal memory 221 and / or the instructions stored in the memory provided in the processor.

[0106] Figure 2B It is a software structure block diagram of the electronic device 200 in the embodiments of the present application. The software system of the electronic device 200 can adopt a layered architecture. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. Taking the Android system as an example, in some embodiments, the Android system can be divided into four layers, from top to bottom are the application layer, application framework layer, Android runtime, and system library, and the kernel layer.

[0107] As Figure 2B shown, the application layer can include application programs such as a gallery service module, a search module, a multimodal understanding module, a natural language understanding module, and a camera application.

[0108] The application framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer includes some predefined functions.

[0109] The system library can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (e.g., OpenGL ES), 2D graphics engine (e.g., SGL), etc. Among them, the media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0110] The kernel layer is the layer between hardware and software.

[0111] Next, in combination with the capture and photo-taking scenario, the working processes of the software and hardware of the electronic device 200 will be exemplarily described.

[0112] When the touch sensor 280K receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including information such as touch coordinates and the timestamp of the touch operation). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking the touch operation as a touch click operation and the control corresponding to the click operation being the control of the camera application icon as an example, the camera application calls the interface of the application framework layer to start the camera application, and then starts the camera driver by calling the kernel layer to capture a static image or video through the camera 293.

[0113] Taking the electronic device as a mobile phone as an example, the visual media search method provided by the embodiments of the present application will be introduced. The visual media search method provided by the embodiments of the present application can be applied in application programs such as the gallery application and the file management application.

[0114] Next, the interfaces and search logics involved in the visual media search method provided by the embodiments of the present application will be exemplarily described in combination with the accompanying drawings.

[0115] As Figure 3A shown in (a) of [reference], the interface 301 displays search history 303 and a "Clear" option 304. The search history 303 includes search statements that the user has entered, such as: "Watching the sunrise on the mountain top", "The sky taken in the photo". Other content displayed on the interface 301 can refer to the relevant display content of the above interface 105, which will not be elaborated here. Among them, the interface 301 can be called a search interface. The mobile phone can respond to the user's operation onFigure 1A Performing a click operation on the search box 104 in the interface 103 shown in (b) in

[0116] As Figure 3A shown in (b) in Figure 3A the interface 305 (i.e., the search result interface), the user enters the search statement "warming the stove and boiling tea" in the search box 306 of the interface 305. The mobile phone searches for 239 pictures. The mobile phone displays some search results of the search statement "warming the stove and boiling tea" (e.g., thumbnails of 8 pictures) and the "More" option 307 corresponding to the search results of the search statement "warming the stove and boiling tea" on the interface 305. In response to the user's operation on the "More" option 307, the mobile phone displays the interface 308 shown in (c) in

[0117] In addition, the mobile phone can also be provided with a negative first screen, a pull-down search interface, etc. It can be understood that the negative first screen can be the leftmost split screen of the electronic device, which is used to provide functions such as search and quick services for the user. Among them, the negative first screen can also be used to display notification messages to be pushed to the user, such as application messages subscribed by the user, real-time hot search messages, segment selection, itinerary information, etc. The pull-down search interface is an interface displayed in response to the user's pull-down operation on the main interface. This interface is used to provide functions such as search and application suggestions for the user. This interface and the interface 315 shown in (c) in Figure 3B below can be the same interface.

[0118] Next, an example will be given with the negative first screen. When the user needs to view the negative first screen of the mobile phone, the user can slide the screen of the mobile phone to make the electronic device display the negative first screen.

[0119] Exemplarily, referring to Figure 3B shown in (a) in Figure 3B the mobile phone can receive the operation 1 performed by the user on the interface 309 (which can be called the desktop) of the mobile phone. Exemplarily, this operation 1 can be a rightward sliding operation shown in (a) in Figure 3B the mobile phone can display the negative first screen 310 shown in (b) in

[0120] The mobile phone receives a click operation by the user on the search box 311 on the negative first screen 310. In response to this click operation, the mobile phone can display the interface 315 as shown in Figure 3B Figure (c) below. The interface 315 may include: a search box 316 and application suggestions. The application suggestions include: icons of applications recommended for use. The interface 315 may also include: a search history 317 and its corresponding "Clear" option 318. In response to a triggering operation by the user on the "Clear" option 318, the search history 317 and the "Clear" option 318 are no longer displayed on the interface 315. In addition, a hot news title, such as "Marathon Race", may be displayed in the search box 316.

[0121] As Figure 3B shown in Figure (d) below, in the search box of the interface 319, a search statement entered by the user, "Peaks photographed on the weekend", is displayed. The interface 319 also displays a preview area 322 of the search results corresponding to the search statement "Peaks photographed on the weekend" by the gallery application and a "Search in Application" option 323 corresponding to the gallery application. In response to a triggering operation by the user on the preview area 322, the mobile phone can enter a photo details interface provided by the gallery application for the user to flip through the search results corresponding to the search statement "Peaks photographed on the weekend", visual media matching the search statement, such as photos or videos. In response to a triggering operation by the user on the "Search in Application" option 323, the mobile phone displays the interface 324 provided by the gallery application as shown in Figure 3B Figure (e) below. The interface 324 displays some search results of the search statement "Peaks photographed on the weekend" and a "More" option corresponding to the search results of the search statement "Peaks photographed on the weekend". In response to a triggering operation by the user on the "More" option, the mobile phone can display a search result details interface, and the search result details interface shows the pictures in the search results of the search statement "Peaks photographed on the weekend". In addition, an online search option 321 may also be displayed in the interface 319. In response to a triggering operation by the user on the online search option 321, the mobile phone displays a search web page and shows online search results in the search web page.

[0122] The interfaces and search logics involved in the visual media search method are introduced above. Next, the specific implementation process of the visual media search method will be continued in combination with the software structure shown in the above Figure 2B figure. As shown in Figure 4 the figure below, the implementation process may include S401 - S423. Among them, S401 - S407 may belong to the index construction stage, and S408 - S423 may belong to the search stage.

[0123] S401. The gallery service module adds and / or modifies visual media and their attributes.

[0124] The attributes of visual media may include one or more of the collection location, collection time, visual media name, face ID of the person in the visual media, person name, and the relationship between the person and the mobile phone user. Taking the captured video or picture as an example, the collection location refers to the shooting location, and the collection time refers to the shooting time; taking the screenshot as an example, the collection location refers to the screenshot location, and the collection time refers to the screenshot time; taking the downloaded video or picture as an example, the collection location refers to the download location, and the collection time refers to the download time.

[0125] The user can add visual media by means of shooting, downloading, screenshotting, etc. In addition, the user can also modify the existing visual media to obtain new visual media. The modifications include but are not limited to operations such as beautification, custom naming, and adding watermarks.

[0126] S402. The gallery service module stores the visual media and its attributes.

[0127] The gallery service module can, in response to the addition or modification operation of the visual media input by the user, store the visual media and its attributes locally on the mobile phone. In practical applications, with the authorization of the user, the mobile phone can store the locally stored visual media and its attributes in the cloud to relieve the storage pressure on the local mobile phone.

[0128] S403. The gallery service module sends Request 1 to the multimodal understanding module. Among them, Request 1 is used to trigger the visual semantic understanding of the visual media.

[0129] S404. The multimodal understanding module returns the image feature vector of the visual media to the gallery service module.

[0130] In the embodiment of the present application, the multimodal understanding module, in response to the above Request 1, determines the image feature vector of the visual media stored in the mobile phone. Since visual semantic understanding requires a large amount of computing resources, in order not to affect the user's use, when the mobile phone is in the state of being charged and the screen is off, the gallery service module can request the multimodal understanding module to perform visual semantic understanding on the visual media stored locally on the mobile phone, such as performing visual semantic understanding on the newly added or modified visual media to obtain the visual semantic vector (or called image feature vector) of the visual media, so as to realize the offline processing of the visual media and reduce the impact of the visual semantic understanding of the visual media on other services running on the mobile phone.

[0131] Among them, the multimodal understanding module can perform visual semantic understanding on visual media based on a multimodal model to obtain the image feature vector of the visual media. In addition, the multimodal model can be used not only for: performing visual semantic understanding on visual media to obtain the visual semantic vector of the visual media; but also for: performing semantic understanding on the search statement to obtain the text feature vector (or called sentence semantic vector) of the search statement. Among them, the process of the multimodal model performing semantic understanding on the search statement can refer to the relevant description below and will not be introduced in detail here.

[0132] In some embodiments, the multimodal model can map visual media and text into vectors of the same dimension. That is to say, the dimension of the visual semantic vector of visual media is the same as that of the semantic vector of text (for example: the sentence semantic vector of the search statement). The multimodal model can specifically be a contrastive language-image pre-training (CLIP) model. The mobile phone can map visual media and text into a unified vector space through the CLIP model to understand the relationship between different modal resources visually and textually, and then use it for image retrieval.

[0133] Among them, the CLIP model is a standard CLIP model, that is, the existing CLIP model, or the CLIP model is a custom CLIP model.

[0134] Exemplarily, the above-mentioned custom CLIP model can include a base model and a fine-tuning model. Among them, the fine-tuning model is obtained by continuing to train the base model through a training set of a specific scenario on the basis of the base model. Therefore, the fine-tuning model has a support scope and can support the processing of images and their texts in a specific scenario. The generalization ability of the fine-tuning model is less than that of the base model, while the retrieval ability of the base model in a specific scenario is less than that of the fine-tuning model.

[0135] In some embodiments, the above-mentioned base model and fine-tuning model reuse part of the network to reduce the resource occupancy of the mobile phone installed with the custom CLIP model. The base model and the fine-tuning model need to process not only visual media but also text. The module for processing visual media in the base model can be called the visual encoding module, and the module for processing visual media in the fine-tuning model can be called the visual fine-tuning module. The above-mentioned image feature vectors can include the image feature vectors corresponding to the base model and the image feature vectors corresponding to the fine-tuning model. The visual encoding module of the base model and the visual fine-tuning module of the fine-tuning model can reuse the image encoder, so that the visual encoding module can use the feature vectors of the visual media output by the image encoder to continue to determine the first L2 norm of the visual media, and realize the determination of the image feature vectors corresponding to the base model. The visual fine-tuning module can use the feature vectors of the visual media output by the image encoder to continue to determine the second L2 norm of the visual media, and realize the determination of the image feature vectors corresponding to the fine-tuning model.

[0136] Exemplarily, after receiving Request 1 sent by the gallery service module, in response to this Request 1, as Figure 5A shown, the multimodal understanding module can input k visual media in the mobile phone into the visual encoding module. The image encoder in the visual encoding module encodes each of the k visual media to obtain the image feature vectors 1 of each visual media. The image feature vector of this visual media can be 768-dimensional, that is, X = {x1, x2,..., x768}, and X represents the image feature vector 1. After that, on the one hand, the visual encoding module can output the 768-dimensional image feature vector 1 so that the image feature vector 1 continues to be used as the input parameter of the visual fine-tuning module. On the other hand, the visual encoding module continues to use the mapping matrix 1 and the image feature vector 1 to calculate and output the first L2 norm α1 of the image feature vectors of each of the k visual media. Among them, the dimension of the mapping matrix 1 is 768*512-dimensional to map the image feature vector 1 from 768 dimensions to 512 dimensions. Here, the image feature vector 1 of the visual media and the first L2 norm of the visual media can be regarded as the image feature vectors of the visual media corresponding to the base model.

[0137] Specifically, for each of the k visual media, the visual encoding module can use to calculate the first L2 norm α1 of the image feature vector of the visual media. Among them, X represents the 768-dimensional image feature vector 1, M 1 represents the mapping matrix 1, V 1 represents the image feature vector 2, and the dimension of V 1 is 512-dimensional, represents the i-th element in V 1 .

[0138] As Figure 5BAs shown, after the visual fine-tuning module receives the image feature vector 1 of each of the k visual media, for each visual media, it continues to use the mapping matrix 2 and the image feature vector 1 of this visual media to calculate and output the second L2 norm α2 of the image feature vector of this visual media. Among them, the dimension of the mapping matrix 2 is 768 * 512, so as to map the image feature vector 1 from 768 dimensions to 512 dimensions. Here, the image feature vector 1 of the visual media and the second L2 norm of the visual media can be regarded as the image feature vector of the visual media corresponding to the fine-tuning model.

[0139] Specifically, for each of the k visual media, the visual encoding module can adopt to calculate the second L2 norm α2 of the image feature vector of the visual media. Among them, X represents the 768-dimensional image feature vector 1, M 2 represents the mapping matrix 2, V 2 represents the image feature vector 3, and the dimension of V 2 is 512 dimensions, represents the i-th element in V 2 .

[0140] In some embodiments, the above fine-tuning model (such as the mapping matrix 2 in the above visual fine-tuning module) is obtained by continuing to train using a training set of a specific scenario. When the device trains the fine-tuning model, it can freeze the image encoder and only train the last fully connected layer M 2 of the fine-tuning model. The training set of this specific scenario can be an image training set including visual content corresponding to a specific whitelist. This specific whitelist represents the support scope of the fine-tuning model. Among them, the device can be the above mobile phone or not. The present application does not limit the device for training the fine-tuning model.

[0141] After obtaining the image feature vector of the visual media (such as the image feature vector of the visual media corresponding to the above base model, the image feature vector of the visual media corresponding to the fine-tuning model), in order to facilitate retrieving the visual media required by the user using the image feature vector of the visual media, the multimedia understanding module can save the image feature vector of the visual media, that is, the above image feature vector 1, the first L2 norm α1 and the second L2 norm α2, so that when saving the image feature vector of a visual media, only a 770-dimensional visual vector needs to be saved, instead of saving 512 * 2 dimensions, that is, 1024-dimensional visual features, reducing the resources required for storing the image feature vector. Among them, the 512-dimensional vector is the image feature vector obtained by processing the 768-dimensional vector using the mapping matrix 1 or the mapping matrix 2 (that is, the above V 1 and V 2 ).

[0142] It should be noted that the above image encoder located in the visual encoding module is only an example. The image encoder can also be located in the visual fine-tuning module, and the present application does not limit it. Additionally, the image feature vectors of the visual media above, including the image feature vectors of the visual media corresponding to the base model and the image feature vectors of the visual media corresponding to the fine-tuning model, are only an example. The image feature vectors of the visual media can also only include the image feature vectors determined by one model, and this model can be a clip model or not a clip model.

[0143] As described above, the determination of the image feature vectors of the visual media in the mobile phone can be determined in an offline state, while the text feature vectors of the search statement can be determined online after the mobile phone receives the search statement input by the user, so that the mobile phone can use the text feature vectors of the search statement and the image feature vectors of the visual media in the mobile phone to determine the visual media that matches the search statement. Among them, the process of determining the text feature vectors of the search statement and using the text feature vectors of the search statement and the image feature vectors of the visual media to determine the visual media that matches the search statement can be referred to below, and will not be introduced in detail here. Next, the process of establishing the index of the visual media will be introduced first.

[0144] S405. The gallery service module stores the image feature vectors of the visual media.

[0145] S406. The gallery service module sends the attribute information and visual semantic vectors of the visual media to the search module.

[0146] In the embodiment of the present application, the gallery service module can store the image feature vectors of each of the k visual media locally in the mobile phone. And the gallery service module can batch send the attribute information and image feature vectors of the visual media to the search module for the search module to construct the index of the visual media.

[0147] Optionally, the multi-gallery service module may not upload the image feature vectors of the visual media to the cloud. Of course, it can also upload the image feature vectors of the visual media to the cloud under user authorization, and the present application does not limit it.

[0148] S407. The search module constructs the index of the visual media.

[0149] The index of the visual media constructed by the search module may include: the attributes of the visual media and / or the visual semantic vectors of the visual media.

[0150] S408. The gallery service module receives the search statement input by the user.

[0151] The user can pass the search interface provided by the gallery service module. For example, when the user is in the above Figure 3AEnter "warming the tea around the stove" in the search box 306 on the interface 305 shown in (b) of []. This "warming the tea around the stove" is the search statement.

[0152] S409. The gallery service module sends the search statement to the search module.

[0153] S410. The search module determines whether the search statement includes visual content.

[0154] In the embodiment of the present application, the search module can determine whether the search statement includes visual content, that is, determine whether it is necessary to search for visual media using the image feature vector and the text feature vector. In other words, the search module determines whether it is necessary to use the multimodal model to determine the visual media.

[0155] In the case where the search statement does not include visual content, it indicates that the visual media can be searched using the attributes of the visual media, without the need to search for the visual media using the image feature vector and the text feature vector, that is, it indicates that there is no need to use the multimodal model to determine the visual media, and the search module can execute S411.

[0156] In the case where the search statement includes visual content, it indicates that it is necessary to search for visual media using the image feature vector and the text feature vector, that is, it indicates that it is necessary to use the multimodal model to determine the visual media, and the search module can execute S412.

[0157] In some embodiments, the search module can determine whether a search statement includes visual content by determining whether the search statement includes a visual semantic entity. The search module can send Request 2 to the natural language understanding module. The natural language understanding module can perform semantic entity recognition on the search statement to obtain the semantic entities included in the search statement. Among them, the natural language understanding module performs semantic entity recognition based on a natural language understanding model. Specifically, the natural language understanding module can use named entity recognition technology (NER) to perform semantic entity recognition on the search statement to obtain the semantic entities included in the search statement. In the embodiments of the present application, the semantic entity can also be referred to as an entity, and the semantic entity can include one or more of: semantic entities related to time, semantic entities related to location, semantic entities related to personal names, semantic entities related to personal relationships, and semantic entities related to visual content (or referred to as visual semantic entities). Among them, optionally, the semantic entity related to visual content is determined from M (M≥1) preset semantic entities related to visual content through named entity recognition technology. These M semantic entities can be configured by developers of the gallery service module according to actual situations. Generally, these M semantic entities are all nouns. Exemplarily, for the search statement "the sky photographed in City 1 on September 1st", semantic entity recognition can obtain three semantic entities: "September 1st", "City 1", and "the sky". Among them, "the sky" is related to visual content and can be a visual semantic entity.

[0158] After that, the natural language understanding module can return the recognized semantic entities to the search module. After that, the search module can determine whether the semantic entities include visual semantic entities. In the case where the semantic entities do not include visual semantic entities, the search module can determine that the search statement does not include visual content. In the case where the semantic entities include visual semantic entities, the search module can determine that the search statement includes visual content.

[0159] S411. The search module performs recall based on the index of the visual media and the search statement to obtain search results.

[0160] Exemplarily, in the case where the search statement does not include a visual semantic entity, it indicates that the search statement does not include content related to visual semantics. The search module can directly query the visual media corresponding to the semantic entity based on the constructed index, that is, based on the attributes of each visual media, and use it as the search result. Here, the semantic entity represents the attributes of the visual media. For example, if the search statement is "September 1st" without visual content, the search module can, based on the index, find the visual media with the acquisition time of September 1st to obtain the search result. Another example is that the search statement includes "photos of Zhang San" without visual content. The search module can match the names of each visual media with "Zhang San" to determine the visual media with the name attribute of "Zhang San" and use it as the search result.

[0161] S412. The search module filters out the non-visual semantic entities in the above search statement to obtain a filtered search statement.

[0162] Exemplarily, in the case where the above semantic entity includes a visual semantic entity, it indicates that the search statement includes content related to visual semantics and may include content unrelated to visual semantics, that is, non-visual semantic entities. Since non-visual semantic entities are irrelevant to the search for visual content, the search module can first filter out the non-visual semantic entities in the search statement to obtain a filtered search statement for searching for visual media that matches the filtered search statement.

[0163] Among them, optionally, the non-visual semantic entity refers to: semantic entities related to time, semantic entities related to location, semantic entities related to personal names, semantic entities related to personal relationships, etc., which are semantic entities related to the attributes of visual media. It should be understood that since personal relationships have been regarded as attributes of visual media before, personal relationships can be regarded as non-semantic entities here.

[0164] For example, for the search statement "the sky photographed this year", "this year" is a non-visual semantic entity. Correspondingly, the filtered search statement is "the sky photographed".

[0165] In practical applications, after filtering out the semantic entities irrelevant to visual content, there may be some redundant stop words. For example, for the search statement "the sky photographed in City 1 this year", after deleting "this year" and "City 1", the stop word "in" becomes a redundant word, so the search module can also filter it out. Specifically, the search module can filter out the non-visual semantic entities irrelevant to visual content and their related stop words in the search statement to obtain a filtered search statement. For example, the filtered search statement corresponding to the search statement "the sky photographed in City 1 this year" is "the sky photographed".

[0166] S413. The search module sends the filtered search statement to the multimodal understanding module.

[0167] S414. The multimodal understanding module performs semantic understanding on the filtered search statement to obtain the text feature vector of the filtered search statement.

[0168] Exemplarily, the multimodal understanding module may use a multimodal model to determine the text feature vector of the filtered search statement. Optionally, the multimodal model may be a CLIP model. The CLIP model may be a standard CLIP model, or the CLIP module is a custom CLIP model.

[0169] In some embodiments, from the above, it can be known that the above image feature vector may include the image feature vector corresponding to the base model and the image feature vector corresponding to the fine-tuning model. The visual encoding module of the base model and the visual fine-tuning module of the fine-tuning model may reuse the image encoder. Correspondingly, in order to keep the similarity of the image-text pair consistent, the present application introduces a text encoding module, so that the text encoding module can reuse the text encoder to respectively output the text feature vector of the filtered search statement corresponding to the base model and the text feature vector of the filtered search statement corresponding to the fine-tuning model, for calculating the similarity between the search statement and the visual media by using the text feature vector of the filtered search statement corresponding to the base model and the image feature vector of the visual media corresponding to the base model, and calculating the similarity between the search statement and the visual media by using the text feature vector of the filtered search statement corresponding to the fine-tuning model and the image feature vector of the visual media corresponding to the fine-tuning model, to realize the calculation of the similarity of the image pair.

[0170] Exemplarily, as Figure 5C shown, the multimodal understanding model inputs the filtered search statement into the CLIP model. The text encoder in the text encoding module of the CLIP model encodes the filtered search statement to obtain and output the feature vector of the filtered search statement, and the dimension of this feature vector is 768 dimensions. Then, the text encoding module may input the feature vector of the filtered search statement into the base model branch and the fine-tuning model branch respectively. Among them, the structures of the base model branch and the fine-tuning model branch are similar.

[0171] For the base model branch: The text encoding module uses the mapping matrix 3 and the feature vector of the filtered search statement to obtain the intermediate variable 1 with 512 dimensions. Exemplarily, the text encoding module calculates the intermediate variable 1 according to 1 T = YN 1 . Where, 1 T 1 represents the intermediate variable 1, Y represents the feature vector of the filtered search statement, and N 1 represents the mapping matrix 3. The text encoding module calculates the first L2 norm of the filtered search statement by using the intermediate variable 1. Exemplarily, the text encoding module may calculate the first L2 norm of the filtered search statement according to Represents the i-th element in T1.

[0172] Since the dimensionality of the image feature vector of the visual media is 768 dimensions, and in order to perform calculations between the image feature vector and the text feature vector, while the intermediate variable 1 of the filtered search statement is 512 dimensions, therefore, the text encoding module needs to map the intermediate variable 1 of the filtered search statement to 768 dimensions. Then, the text encoding module can use the above mapping matrix 1 and the first L2 norm of the filtered search statement to map the intermediate variable 1 into a 768-dimensional text feature vector to obtain the text feature vector of the filtered search statement corresponding to the base model. Specifically, the text encoding module can use To calculate the text feature vector of the filtered search statement corresponding to the base model. Where, T' 1 Represents the text feature vector of the filtered search statement corresponding to the base model Is M 1 The transpose matrix of, M 1 Represents the above mapping matrix 1.

[0173] For the fine-tuning model branch: The text encoding module uses the mapping matrix 4 and the feature vector of the filtered search statement to obtain the intermediate variable 2 of 512 dimensions. Exemplarily, the text encoding module calculates the intermediate variable 2 according to T 2 = YN 2 , Calculate the intermediate variable 2. Where, T 2 Represents the intermediate variable 2, N 2 Represents the mapping matrix 4, and the text encoding module calculates the second L2 norm of the filtered search statement using the intermediate variable 2. Exemplarily, the text encoding module can calculate according to To calculate the second L2 norm of the filtered search statement. Where, β 2 Represents the second L2 norm of the filtered search statement, Represents T 2 The i-th element in.

[0174] Since the dimensionality of the image feature vector of the visual media is 768 dimensions, and in order to perform calculations between the image feature vector and the text feature vector, while the intermediate variable 2 of the filtered search statement is 512 dimensions, therefore, the text encoding module needs to map the intermediate variable 2 of the filtered search statement to 768 dimensions. Then, the text encoding module can use the above mapping matrix 2 and the first L2 norm of the filtered search statement to map the intermediate variable 2 into a 768-dimensional text feature vector to obtain the text feature vector of the filtered search statement corresponding to the fine-tuning model. Specifically, the text encoding module can use To calculate the text feature vector of the filtered search statement corresponding to the base model. Where, T' 2 Represents the text feature vector of the filtered search statement corresponding to the fine-tuning model, Is M 2The transposed matrix of, M 1 represents the above mapping matrix 2.

[0175] It should be noted that the text feature vector of the filtered search statement corresponding to the fine-tuning model and the text feature vector of the filtered search statement corresponding to the base model can be output by the text encoding module simultaneously.

[0176] In the embodiments of the present application, the base model and the fine-tuning model in the custom CLIP model reuse the text encoder, so that the custom CLIP model only needs to calculate the feature vector of the filtered search statement once, and then can use the feature vector of the filtered search statement to respectively determine the text feature vector of the filtered search statement corresponding to the base model and the text feature vector of the filtered search statement corresponding to the fine-tuning model, without the base model and the fine-tuning model calculating the feature vector of the filtered search statement separately, thereby improving the calculation efficiency of the text feature vector of the filtered search statement.

[0177] In the embodiments of the present application, in order to ensure the generalization ability of the multi-modal model and the ability to recognize specific images and texts, the present application sets the multi-modal model to include a base model and a fine-tuning model. Since the base model and the fine-tuning model have the same model weights, that is, part of the network is the same, directly setting two models will cause waste of mobile phone memory. Therefore, the present application reuses the image encoder and the text encoder for the base model and the fine-tuning model, so that the base model and the fine-tuning model can respectively use the feature vectors output by the image encoder and the text encoder to determine their respective corresponding text feature vectors and image feature vectors. Generally speaking, the determination time of the image feature vector and the text feature vector is reduced by nearly half.

[0178] S415. The multi-modal understanding module performs recall based on the text feature vector and the image feature vector of the visual media to obtain candidate visual media.

[0179] In some embodiments, the image feature vector of the above visual media can be sent by the search module to the multi-modal understanding module, or can be obtained by the multi-modal understanding module from the local of the mobile phone.

[0180] In the embodiments of the present application, for each visual media on the mobile phone (such as each visual media in the above k visual media), the multi-modal understanding module can calculate the vector similarity (or simply referred to as similarity) between the image feature vector of the visual media and the text feature vector of the filtered search statement, that is, calculate the similarity between the visual media and the filtered search statement based on the image feature vector of the visual media and the text feature vector of the filtered search statement. Then, the multi-modal understanding module determines the visual media that matches the filtered search statement from the k visual media according to the vector similarity, and uses the determined visual media as the candidate visual media.

[0181] Exemplarily, the multimodal understanding module may determine visual media with a vector similarity greater than or equal to threshold 1 as visual media matching the filtered search statement. Optionally, when the number of visual media with a vector similarity greater than or equal to threshold 1 is greater than quantity 1, the multimodal understanding module may sort the visual media with a vector similarity greater than or equal to threshold 1 in descending order of vector similarity. Subsequently, the multimodal understanding module may determine the top n visual media as visual media matching the filtered search statement, where n is a positive integer.

[0182] In some embodiments, the image feature vector of the above-mentioned visual media may include the image feature vector of the visual media corresponding to the base model and the image feature vector of the visual media corresponding to the fine-tuning model. The text feature vector of the above-mentioned filtered search statement includes the text feature vector of the filtered search statement corresponding to the base model and the text feature vector of the filtered search statement corresponding to the fine-tuning model. In one case, as Figure 5D shown, the multimodal understanding module may calculate the vector similarity 1 between the text feature vector of the filtered search statement corresponding to the base model and the image feature vector of the visual media corresponding to the base model (i.e., calculate the similarity between the visual media corresponding to the base model and the filtered search statement), and calculate the vector similarity 2 between the text feature vector of the filtered search statement corresponding to the fine-tuning model and the image feature vector of the visual media corresponding to the fine-tuning model (i.e., calculate the similarity between the visual media corresponding to the fine-tuning model and the filtered search statement). Subsequently, the multimodal understanding module may determine visual media with vector similarity 1 or vector similarity 2 greater than or equal to threshold 1 as visual media matching the filtered search statement.

[0183] In another case, the multimodal understanding module may use whether the filtered search statement hits the whitelist 1 to determine whether the filtered search statement is within the support range of the fine-tuning model, that is, determine whether to use the fine-tuning model branch to determine candidate visual media. The process of the multimodal understanding module using the whitelist 1 to determine candidate visual media will be described below in combination with Figure 6 , to introduce the process of the multimodal understanding module using the whitelist 1 to determine candidate visual media.

[0184] S501. The multimodal understanding module determines whether the filtered search statement belongs to the whitelist 1.

[0185] In the embodiments of the present application, when the filtered search statement does not belong to the whitelist, it indicates that the filtered search statement is within the support range of the base model. The multimodal understanding module may use the base model branch to determine candidate visual media, and the multimodal understanding module may execute S502.

[0186] In the case where the filtered search statement belongs to the whitelist, it indicates that the search statement is within the support range of the fine-tuning model. The multimodal understanding module can enable the fine-tuning model branch to determine candidate visual media, and the multimodal understanding module can execute S504.

[0187] In some embodiments, the multimodal understanding module determines whether each semantic entity (i.e., visual semantic entity) in the filtered search statement belongs to Whitelist 1. Considering that the search statement input by the user is generally a phrase, the multimodal understanding module can determine that the filtered search statement belongs to the whitelist when all semantic entities in the filtered search statement belong to Whitelist 1. In the case where there is a semantic entity that does not belong to Whitelist 1, it is determined that the filtered search statement does not belong to the whitelist. For example, the filtered search statement is "a boy holding a basket", and the semantic entities include "basket" and "boy". The multimodal understanding module can respectively determine whether the basket and the boy belong to Whitelist 1. When both the basket and the boy belong to Whitelist 1, the multimodal understanding module can determine that the filtered search statement belongs to the whitelist. In the case where the basket or the boy does not belong to Whitelist 1, the multimodal understanding module can determine that the filtered search statement does not belong to the whitelist.

[0188] S502. For each visual media, the multimodal understanding module calculates the vector similarity 1 between the image feature vector of the visual media corresponding to the base model and the text feature vector of the filtered search statement corresponding to the base model.

[0189] In the embodiments of the present application, as Figure 7 shown, in the case where the filtered search statement does not belong to Whitelist 1, the multimodal understanding module can call the text feature vector of the filtered search statement output by the base model branch of the text encoding module and the image feature vector of the visual media corresponding to the base model, and calculate the score corresponding to the visual media corresponding to the base model, that is, calculate the vector similarity 1 between the visual media corresponding to the base model and the filtered search statement.

[0190] Specifically, the multimodal understanding module can calculate the score corresponding to the visual media corresponding to the base model through where S 1 represents the score corresponding to the visual media corresponding to the base model, α1 represents the first L2 norm of the image feature vector of the above visual media. X represents the above image feature vector 1, and T' 1 represents the text feature vector of the filtered search statement corresponding to the above base model.

[0191] S503. The multimodal understanding module uses the visual media with a vector similarity 1 greater than the threshold 1 as the candidate visual media.

[0192] S504. For each visual medium, the multimodal understanding module calculates the vector similarity 2 between the image feature vector of the visual medium corresponding to the fine-tuning model and the text feature vector of the filtered search statement corresponding to the fine-tuning model.

[0193] In the embodiments of the present application, as described above Figure 7 shown, when the filtered search statement belongs to the whitelist 1, the multimodal understanding module can call the text feature vector of the filtered search statement output by the fine-tuning model branch of the text encoding module and the image feature vector of the visual medium corresponding to the fine-tuning model, and calculate the score corresponding to the visual medium corresponding to the fine-tuning model, that is, calculate the vector similarity 2 between the visual medium corresponding to the fine-tuning model and the filtered search statement.

[0194] Specifically, the multimodal understanding module can pass through Calculate the score corresponding to the visual medium corresponding to the base model. Where S 2 represents the score corresponding to the visual medium corresponding to the fine-tuning model, and α 2 represents the second L2 norm of the image feature vector of the above visual medium. X represents the above image feature vector 1, and T' 2 represents the text feature vector of the filtered search statement corresponding to the above fine-tuning model.

[0195] S505. The multimodal understanding module uses the visual media with a vector similarity 2 greater than the threshold 1 as candidate visual media.

[0196] In some embodiments, the operations performed by the above multimodal understanding module can be performed by a multimodal model. In addition, the above S501-S505 can be performed by the output module of the model in the above multimodal understanding module, that is, the multimodal model.

[0197] In the embodiments of the present application, the multimodal understanding module can determine the image feature vector of the visual medium offline, and only needs to determine the text feature vector of the search statement online, which improves the calculation time of the vector similarity between the image feature vector and the text feature vector, thereby effectively reducing the user's retrieval time. In addition, the multimodal understanding module can judge whether to use the fine-tuning model to determine the candidate visual media according to whether the filtered search statement belongs to the whitelist 1, so as to improve the determination accuracy of the candidate visual media, thereby improving the accuracy of the search results.

[0198] Optionally, the above whitelist 1 can be extended, that is to say, the support range of the fine-tuning model can be extended. In other words, the scenarios supported by the fine-tuning model can be extended. It only needs to train the fine-tuning model with the training set corresponding to the extended scenario. It should be understood that training the fine-tuning model is actually training the fine-tuning model branch in the above fine-tuning module and text encoding module, without training the reused parts (such as the above image encoder, text encoder).

[0199] In some embodiments, after determining the candidate visual media, the multimodal understanding module may directly use the above candidate visual media as the search result. After that, the multimodal understanding module may return the search result to the search module. Then, the search module may send the search result to the gallery service module for the gallery service module to display the search result.

[0200] In addition, to improve the accuracy of the search result, after obtaining the above candidate visual media, the mobile phone may continue to perform a secondary confirmation on the candidate visual media to further screen visual media from the candidate visual media to obtain visual media that matches the filtered search statement, i.e., the search statement. Optionally, the mobile phone may use both visual media with vector similarity 1 greater than threshold 1 and visual media with vector similarity 2 greater than threshold 1 as candidate visual media. In other words, the mobile phone may use the visual media determined by the base model and the fine-tuned model respectively as candidate visual media, i.e., perform two searches. The process of secondary confirmation of the candidate visual media is described below.

[0201] S416. The multimodal understanding module sends the information of the filtered search statement and the image information of the candidate visual media to the secondary confirmation module.

[0202] The information of the filtered search statement includes one or more of the filtered search statement, the word segmentation result of the filtered search statement, label 1 included in the filtered search statement, and the text feature vector of the filtered search statement.

[0203] The image information of the candidate visual media includes one or more of the similarity between the candidate visual media and the filtered search statement, label 2 of the candidate visual media, and the image feature vector of the candidate visual media.

[0204] In some embodiments, the above image feature vector and text feature vector are determined based on a multimodal model including a base model and a fine-tuned model. Correspondingly, the text feature vector of the filtered search statement may include the text feature vector of the filtered search statement corresponding to the base model and the text feature vector of the filtered search statement corresponding to the fine-tuned model. The image feature vector of the candidate visual media may include the image feature vector of the candidate visual media corresponding to the base model and the image feature vector of the candidate visual media corresponding to the fine-tuned model.

[0205] The similarity between the above candidate visual media and the filtered search statement may include the similarity between the candidate visual media corresponding to the base model and the filtered search statement (or referred to as similarity 1), and the similarity between the candidate visual media corresponding to the fine-tuned model and the filtered search statement (or referred to as similarity 2).

[0206] Of course, the above-mentioned image feature vectors, text feature vectors, and similarity can also be determined by only one model (such as the above-mentioned base model or fine-tuning model), and the present application does not limit this.

[0207] In some embodiments, the multimodal understanding module can perform word segmentation extraction on the filtered search statement to obtain the word segmentation result of the filtered search statement. For example, if the filtered search statement is "The child is playing by the sea", the word segmentation result is: child, is, by the sea, and playing. Exemplarily, the multimodal understanding module can perform word segmentation extraction on the filtered search statement through the natural language understanding engine service (the natural language understanding, NLU). Optionally, the multimodal understanding module can perform word segmentation extraction on the filtered search statement through the natural language understanding module.

[0208] After obtaining the word segmentation result of the filtered search statement, the multimodal understanding module can perform label mapping on the word segmentation result to obtain the label included in the filtered search statement, that is, label 1. Specifically, for each word in the word segmentation result, the filtered search module determines whether the word exists in the preset label table. In the case where the word does not exist in the preset label table, the multimodal understanding module can determine that the word does not have a corresponding label, that is, the filtered search statement does not include the label corresponding to the word.

[0209] In the case where the word exists in the preset label table, the multimodal understanding module can use the label corresponding to the word in the preset label table as the label 1 included in the filtered search statement. For example, the word segmentation includes "cat", and the label corresponding to "cat" in the preset label table is "cat", so the label included in the filtered search statement includes "cat". It should be noted that the words in the word segmentation result of the filtered search statement may be the same as or different from the labels corresponding to the words.

[0210] In some embodiments, the above-mentioned label 2 of the candidate visual media is obtained from the label library. The label 2 of the visual media in the label library represents the classification label of the visual media, which can be determined by the mobile phone (such as the multimodal understanding module in the mobile phone) using the image classification model to identify the visual media. Exemplarily, the classification label of the visual media can also be determined offline by the mobile phone.

[0211] S417. The secondary confirmation module determines the Boolean value corresponding to each candidate visual media based on the information of the filtered search statement and the image information of the candidate visual media.

[0212] S418. For each candidate visual media, when the Boolean value corresponding to the candidate visual media is true, the secondary confirmation module determines that the candidate visual media is visual media 1.

[0213] In the embodiments of the present application, the multimodal understanding module inputs the information of the filtered search statement and the image information of the candidate visual media into the secondary confirmation module, so that the secondary confirmation module determines the Boolean value corresponding to each candidate visual media, and the Boolean value corresponding to the candidate visual media indicates whether the candidate visual media is a visual media matching the filtered search statement, realizing the secondary confirmation of the candidate visual media.

[0214] In some embodiments, the above-mentioned secondary confirmation module may include a binary classification module, a dynamic threshold module, a label confirmation module, and a whitelist threshold module. The binary classification module may adopt a binary classification model to determine whether there is a correlation between the filtered search statement and the candidate visual media, so as to obtain the Boolean value of the candidate visual media. The dynamic threshold module may adopt a dynamic threshold model to determine the threshold 2 that matches the length of the filtered search statement, so that the similarity between the candidate visual media and the filtered search statement can be compared with the threshold 2 to obtain the Boolean value of the candidate visual media. The label confirmation module may compare the label 2 of the candidate visual media with the label 1 included in the filtered search statement to obtain the Boolean value of the candidate visual media. The whitelist threshold module may determine the threshold 3 that matches the filtered search statement as a whole according to whether the filtered search statement hits the preset dictionary, so that the similarity between the candidate visual media and the filtered search statement can be compared with the threshold 3 to obtain the Boolean value of the candidate visual media. Among them, the detailed process of the secondary determination module determining the candidate visual media through the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module, that is, the implementation process of the above S417, can refer to the relevant descriptions below and will not be introduced here first.

[0215] S419. The secondary confirmation module returns the visual media 1 to the search module.

[0216] S420. The search module obtains the visual media 2 based on the non-visual semantic subject in the above search statement and the index of the visual media.

[0217] S421. The search module takes the intersection of the visual media 1 and the visual media 2 as the search result.

[0218] In the embodiments of the present application, the search module searches for the visual media (or referred to as visual media 2) corresponding to the filtered non-visual semantic subject in the search statement input by the user from the constructed index, that is, searches for the visual media whose attributes match the non-visual semantic subject, and takes it as the visual media 2.

[0219] After that, the search module can determine the intersection of Visual Media 2 and Visual Media 1 to obtain the target visual media. The attributes of the target visual media match the non-visual semantic entity of the search statement, and the visual content of the target search media corresponds to the visual semantic entity of the search statement. For example, if the search statement is "the sky photographed this year", then Visual Media 1 includes sky content, and the acquisition time of Visual Media 2 is this year. Then, the target visual media includes sky content, and the acquisition time of the target visual media is this year.

[0220] S422. The search module sends the search results to the gallery service module.

[0221] S423. The gallery service module displays the search results.

[0222] In the embodiments of the present application, the gallery service module can display the search results to the user. For example, through Interface 305 described above, Figure 3A Interface 308 described above Figure 3A to display the search results. As shown in Interface 305 described above Figure 3A the search results include 239 pictures, and Interface 305 only displays the thumbnails of 8 pictures in the search results; if the user wants to view these 239 pictures, the user can click the "More" option in Interface 305, and in response to this operation, the mobile phone displays Interface 308 as shown in Figure 3A Interface 308 described above.

[0223] In some embodiments, the gallery service module can sort the target visual media in the search results according to the acquisition time and display the target visual media. For example, sort the target visual media in ascending order of the acquisition time, so that the earlier the acquisition time, the higher the display order.

[0224] In other embodiments, the gallery service module can sort the target visual media in the search results according to the similarity between the target visual media and the filtered search statement. For example, sort the target visual media in descending order of the similarity, so that the higher the similarity of the target visual media, the higher the display order, that is, the similarity between the target visual media displayed earlier and the filtered search statement is greater than or equal to the similarity between the target visual media displayed later and the filtered search statement. Here, the similarity between the target visual media and the filtered search statement can be Similarity 1 or Similarity 2, or a similarity calculated based on Similarity 1 and Similarity 2.

[0225] Optionally, when the filtered search statement belongs to the above-mentioned whitelist 1, the similarity between the target visual media and the filtered search statement can be similarity 2, that is, the similarity between the visual media corresponding to the fine-tuning model and the filtered search statement. When the filtered search statement does not belong to the above-mentioned whitelist 1, the similarity between the target visual media and the filtered search statement can be similarity 1, that is, the similarity between the visual media corresponding to the base model and the filtered search statement.

[0226] Optionally, the similarity calculated based on similarity 1 and similarity 2 can be the average of similarity 1 and similarity 2. Or, calculate the weighted sum of similarity 1 and similarity 2, and this application does not limit it.

[0227] Next, a possible implementation process of the above S417 will be continued, as Figure 8 shown, this process may include S417a - S417g.

[0228] S417a. The secondary confirmation module determines whether the filtered search statement exists in the preset fine-tuning model vocabulary.

[0229] In the embodiments of the present application, the secondary confirmation module can determine whether the filtered search statement exists in the preset fine-tuning model vocabulary to determine whether the filtered search statement corresponds to the branch of the base model or the branch of the fine-tuning model, that is, to determine whether to perform secondary confirmation through the module corresponding to the base model or the module corresponding to the fine-tuning model.

[0230] When the filtered search statement does not exist in the above-mentioned preset fine-tuning vocabulary, it indicates that the filtered search statement corresponds to the branch of the base model, that is, it indicates that the module corresponding to the base model is used to perform secondary confirmation on the candidate visual media to select a visual media that better matches the filtered search statement from the candidate visual media, and the secondary confirmation module can execute S417b.

[0231] When the filtered search statement exists in the above-mentioned preset fine-tuning vocabulary, it indicates that the filtered search statement corresponds to the branch of the fine-tuning model, that is, it indicates that the module corresponding to the fine-tuning model is used to perform secondary confirmation on the candidate visual media to select a visual media that better matches the filtered search statement from the candidate visual media, and the secondary confirmation module can execute S417f.

[0232] S417b. The secondary confirmation module inputs the text feature vector of the filtered search statement and the image feature vectors of each candidate visual media into the binary classification module, and obtains the Boolean value 1 corresponding to each candidate visual media output by the binary classification module.

[0233] In the embodiments of the present application, it is assumed that the number of candidate visual media is m. As Figure 9AAs shown, the secondary confirmation module inputs the image feature vectors of each candidate visual media among the m candidate visual media and the text feature vector of the filtered search statement into the binary classification model in the binary classification module. For each candidate visual media among the m candidate visual media, the binary classification model determines the matching degree between the candidate visual media and the search statement based on the image feature vector of the candidate visual media and the text feature vector of the filtered search statement. This matching degree represents the degree of correlation between the candidate visual media and the filtered search statement. The greater the matching degree, the more relevant the candidate visual media is to the filtered search statement. After that, the binary classification model compares the matching degree between the candidate visual media and the filtered search statement with the classification threshold to determine whether the candidate visual media matches the search statement, obtaining the Boolean value (bool) 1 corresponding to the candidate visual media, and thus outputs the Boolean values 1 corresponding to the m candidate visual media. When the matching degree is greater than the classification threshold, it indicates that the candidate visual media matches the filtered search statement, and the Boolean value 1 corresponding to the candidate visual media is true. When the matching degree is less than or equal to the classification threshold, it indicates that the candidate visual media does not match the filtered search statement, and the Boolean value 1 corresponding to the candidate visual media is false.

[0234] Among them, the binary classification model is a model pre-trained using a training sample set, and it can determine whether these two match based on the image feature vector and the text feature vector. This binary classification model can be trained by the above-mentioned mobile phone or other devices. Below, taking the binary classification model trained by a mobile phone as an example, the training process of the binary classification model will be introduced.

[0235] In some embodiments, when the filtered search statement exists in the above-mentioned preset fine-tuning model vocabulary, it indicates that the filtered search statement corresponds to the branch of the fine-tuning model. Since the branch corresponding to the fine-tuning model does not include the binary classification module, the secondary confirmation module does not need to execute S417b above.

[0236] When the word segmentation of the filtered search statement does not exist in the above-mentioned preset fine-tuning model vocabulary, it indicates that the filtered search statement corresponds to the branch of the base model. The text feature vector of the filtered search statement in S417b above may include the text feature vector of the filtered search statement corresponding to the base model, and the image feature vector of the candidate visual media in S417b above may include the image feature vector of the candidate visual media corresponding to the base model.

[0237] It can be understood that if the input parameters of the secondary confirmation module only include one image feature vector of the candidate visual media and one text feature vector of the filtered search statement, and do not include the image feature vector of the candidate visual media corresponding to the fine-tuning model and the image feature vector of the candidate visual media corresponding to the base model, then the image feature vector of the candidate visual media in the above S417b is the input image feature vector of the candidate visual media, and the text feature vector of the filtered search statement is the input text feature vector of the filtered search statement.

[0238] Exemplarily, the training sample set of the above binary classification model includes multiple pieces of training data, and each piece of training data in the multiple pieces of training data includes a sample image and its corresponding description text. The training sample set may include a positive sample training set (abbreviated as positive samples) and a negative sample training set (abbreviated as negative samples). The sample image in the training data included in the positive samples matches the description text corresponding to the sample image. For example, as Figure 9B shown in the image, the description text corresponding to this image is "a little boy holding a basket". Since Figure 9B the visual content expressed is that the little boy is holding a basket, which matches the corresponding description text, therefore, this image and its corresponding description text can be used as a piece of training data in the positive samples.

[0239] The sample image in the training data included in the negative samples does not match the description text corresponding to the sample image. For example, as Figure 9B shown in the image, the description text corresponding to this image is "flowers". Since Figure 9B the visual content expressed does not match the description text, therefore, this sample image and its corresponding description text can be used as a piece of training data in the negative samples.

[0240] It should be noted that since users generally prefer to use natural language when searching for visual media, the description text corresponding to the sample image also uses natural language, which conforms to the actual search situation.

[0241] To improve the training accuracy, it is necessary to ensure the quantity of the training data in the positive and negative samples. Considering that the efficiency of obtaining positive and negative samples manually is relatively low, the mobile phone can automatically generate positive and negative samples. Below, taking the above binary classification model as a text-image matching model as an example, the generation process of positive and negative samples will be continued. As Figure 10 shown, the process is as follows:

[0242] S1. The binary classification module obtains P sample data pairs. Among them, each sample data pair in the P sample data pairs includes a sample image and its corresponding description text.

[0243] S2. The binary classification module extracts sub-texts in the description text of each sample data.

[0244] Exemplarily, the above-mentioned sub-text may include descriptive phrases and / or nouns in the descriptive text. Exemplarily, the sub-text generally does not include adverbials, verbs, etc. in the descriptive text that do not have corresponding actual visual content. The binary classification module can tokenize the descriptive text to determine the descriptive phrases and nouns in the descriptive text, and obtain the sub-text in the descriptive text. For example, if the descriptive text is "a little boy sitting in a basket", the sub-text in this descriptive text is "basket", "little boy", and "boy". Another example, if the descriptive text is "a little boy sitting in a basket and sucking his thumb". After tokenizing this descriptive text, the obtained sub-text is "basket", "boy", "little boy", "little boy sucking his thumb", where "little boy sucking his thumb" can represent a descriptive phrase.

[0245] In some embodiments, the binary classification module can use the Han Language Processing Package (HanLP) to tokenize the descriptive text corresponding to the sample image.

[0246] S3. For each sample image, the binary classification module generates a text set corresponding to the sample image based on the descriptive text corresponding to the sample image and the sub-text in the descriptive text.

[0247] Among them, the text set corresponding to the sample image represents a description set of the visual content corresponding to the sample image. The text set corresponding to the sample image may include the descriptive text corresponding to the sample image and each sub-text in the descriptive text. For example, the descriptive text corresponding to the sample image is "a little boy holding a basket", and the sub-text in the descriptive text includes "basket", "little boy", "boy". Correspondingly, the text set corresponding to this image is {"basket", "little boy", "boy", "a little boy holding a basket"}.

[0248] S4. The binary classification module randomly selects a text element from the text sets corresponding to each sample image to obtain a text element 1 corresponding to each sample image.

[0249] Among them, the text element in the text set can be a sub-text or a descriptive text.

[0250] In some embodiments, the binary classification module can randomly select text elements from the text sets corresponding to each sample image according to a preset ratio. The preset ratio includes the ratio of the selected text elements that are descriptive texts and the ratio of the selected text elements that are sub-texts. For example, the preset ratio is 8:2, the probability that the selected text element is a descriptive text is 80%, and the probability that the selected text element is a sub-text is 20%. Since generally the image encoder and the text encoder are trained using the image and its corresponding text as a whole, the ratio of the selected text elements that are descriptive texts in the preset ratio will be greater than the ratio of the selected text elements that are sub-texts to ensure the training effect of the model.

[0251] S5. For each sample image, the binary classification module respectively uses the P text elements 1 corresponding to the P sample images as the sample text elements corresponding to this sample image.

[0252] Among them, the number of sample text elements corresponding to one sample image among the P sample images is P. One sample image and one sample text element corresponding to it form a sample, so that multiple samples can be obtained. Since it is necessary to use positive and negative samples to train the model, after obtaining the samples, it is necessary to distinguish the positive and negative samples in the samples, so as to train the binary classification model with the positive and negative samples. The process of distinguishing the positive and negative samples in the samples will be introduced below.

[0253] In the embodiment of the present application, the binary classification module initially obtains P sample data pairs. These P sample data pairs are all positive samples. However, the training of the binary classification model also requires negative samples. Therefore, the binary classification module can perform image-text matching on the P sample images and one text element (i.e., text element 1) in the text set corresponding to the P sample images. That is, for each sample image among the P sample images, the binary classification module can use all P text elements 1 as the sample text elements corresponding to this sample image. This sample image and each of the P text elements 1 form a sample, so that P*P samples can be obtained, realizing an increase in the number of samples and thus realizing the rapid generation of samples. In order to implement the training of the binary classification model, the binary classification model needs to distinguish the positive and negative samples among the P*P samples. The process of distinguishing the positive and negative samples will be continued to be introduced below.

[0254] S6. For each sample text element corresponding to this sample image, the binary classification module determines whether this sample text element belongs to the text set corresponding to this sample image.

[0255] In the embodiment of the present application, for each sample text element corresponding to the sample image, the binary classification module can determine whether this sample text element belongs to the text set corresponding to this sample image, so as to determine whether this sample text element corresponds to the visual content of this sample image, and thus determine whether the sample composed of this sample text element and this sample image is a positive sample.

[0256] In the case where this sample text element is in the text set corresponding to this sample image, it indicates that this sample text element corresponds to (i.e., matches) the visual content of this sample image. That is to say, the sample composed of this sample text element and this sample image is a positive sample. Therefore, the binary classification module can execute S7.

[0257] In the case where this sample text element is not the text set corresponding to this sample image, it indicates that this sample text element does not correspond to (i.e., does not match) the visual content of this image. That is to say, the sample composed of this sample text element and this sample image is not a positive sample. Therefore, the binary classification module can execute S8.

[0258] S7. The binary classification module uses the sample image and the sample text element as positive samples.

[0259] S8. The binary classification module uses the image and the sample text element as negative samples.

[0260] In some embodiments, after obtaining the sample text elements corresponding to each sample image, the binary classification module may generate a corresponding sample matrix. Among them, the i-th row element in the sample matrix represents each sample text element corresponding to the i-th sample image, and the j-th column element in the sample matrix represents the text element 1 corresponding to the j-th sample image, that is, the text element randomly selected from the text set corresponding to the j-th sample image.

[0261] For example, the sample images include image a, image b, image c, and image d. Figure 11A The "1" in the sample matrix shown is the text element 1 corresponding to image a (i.e., the first sample image). And so on, "6" is the text element 1 corresponding to image d (i.e., the fourth sample image). The first row element 50 (i.e., 1, 2, 5, 6) is the sample text element corresponding to image a. The second column element 51 (i.e., 2, 2, 2, 2) is the text element 1 corresponding to image b.

[0262] Among them, a sample text element in the sample matrix and the sample image corresponding to the row where the sample text element is located form a sample. For example, as described above Figure 11A The "1" in the first row and first column shown forms a sample with image a.

[0263] In the embodiments of the present application, the binary classification model generates a corresponding sample matrix by mutually matching P sample images and P text elements 1, so that the binary classification module can distinguish whether the sample text elements in the sample matrix belong to positive samples or negative samples, realize the rapid determination of positive and negative samples, improve the determination efficiency of positive and negative samples, and moreover, by generating a label matrix, the comprehensiveness of image-text matching can be guaranteed, and the situation of missing sample images or text elements 1 can be avoided, such as avoiding the situation that a certain sample image is not matched with a certain text element 1.

[0264] After obtaining the sample matrix, for each sample text element in the sample matrix, the binary classification module needs to determine whether the sample text element matches the sample image corresponding to the row where the sample text element is located. To improve the matching efficiency, the binary classification module can use the text intersection matrix 1 to match with the sample matrix to obtain positive and negative samples. The text intersection matrix 1 represents the intersection of the text sets corresponding to the sample images. Exemplarily, the determination process of the text intersection matrix 1 may include:

[0265] The binary classification module determines the overlapping elements in the text sets corresponding to any two sample images. Subsequently, the binary classification module generates a corresponding text intersection matrix 1. Among them, the s-th element in the t-th row of the text intersection matrix 1 represents the overlapping elements in the text sets between the t-th sample image and the s-th sample image.

[0266] For example, the text set corresponding to image a (as Figure 11B shown) is {1, 2, 3}, the text set corresponding to image b is {2, 3, 4}, the text set corresponding to image c is {5, 6}, and the text set corresponding to image d is {6, 7}. It should be understood that the numbers here are actually corresponding text elements (such as the above-mentioned sub-texts, descriptive texts).

[0267] Subsequently, the binary classification module can determine the overlapping elements (i.e., the same elements) in the text sets corresponding to any two sample images among image a, image b, image c, and image d. For example, the same elements in the text sets corresponding to image a and image b are 2, 3.

[0268] Subsequently, the binary classification module generates a text intersection matrix 1 as Figure 11C shown. This text intersection matrix 1 includes the overlapping elements in the text sets corresponding to the sample image and each sample image (i.e., image a, image b, image c, and image d). Among them, Figure 11C the element in the first row and first column of the text intersection matrix 1 shown represents the overlapping elements in the text set between image a and image a (i.e., the text set corresponding to image a). The element in the first row and second column represents the overlapping elements in the text set between image a and image b. And so on, Figure 11C the element in the fourth row and fourth column in

[0269] In some embodiments, after the binary classification module obtains the text sets corresponding to each sample image among P sample images, it can calculate the union of the text sets corresponding to each sample image. The union of this text set can include the text elements in all text sets. Subsequently, the binary classification module can assign numbers (such as the above-mentioned numbers) to each text element in the union of the text sets, so that different text elements in the text set correspond to different numbers, and the same text elements in different text sets correspond to the same numbers. By assigning numbers to text elements, the graphic-text matching efficiency can be improved, and thus the generation efficiency of positive and negative samples can be improved.

[0270] The determination process of the text intersection matrix 1 is introduced above. Next, the process of using the text intersection matrix 1 to match with the sample matrix to determine positive and negative samples will be continued.

[0271] The binary classification module intersects the above-mentioned text intersection matrix 1 and the sample matrix to obtain text intersection matrix 2.

[0272] After that, for each intersection element in text intersection matrix 2, the binary classification module determines whether the intersection element is empty.

[0273] When the intersection element is not empty, it indicates that the visual content of the intersection element matches the visual content of the sample image corresponding to its row. Therefore, the binary classification module can confirm that the intersection element and the sample image corresponding to its row are positive samples.

[0274] When the intersection element is empty, it indicates that the visual content of the intersection element does not match the visual content of the sample image corresponding to its row. Therefore, the binary classification module can confirm that the element and the sample image corresponding to its row are negative samples, thus realizing the rapid determination of positive and negative samples. Among them, the element corresponding to the sample image can also be called the text corresponding to the sample image.

[0275] Optionally, the binary classification module can use 1 and 0 to distinguish whether the intersection elements in text intersection matrix 2 belong to positive samples or negative samples, that is, to distinguish whether the sample text elements in the sample matrix belong to positive samples or negative samples. When the intersection element in text intersection matrix 2 is empty, the binary classification module can set the label corresponding to the intersection element to 0, that is, label the sample text element corresponding to the intersection element with 0.

[0276] When the intersection element in text intersection matrix 2 is not empty, the binary classification module can set the label corresponding to the element to 1, that is, label the sample text element corresponding to the intersection element with 1, thereby obtaining a positive and negative sample label matrix and realizing the setting of labels for the sample text elements in the sample matrix. The position of the intersection element in text intersection matrix 2 is the same as the position of the sample text element corresponding to the intersection element in the sample matrix.

[0277] Based on this, the binary classification module can use the sample text element corresponding to the label 0 and the image corresponding to the row where the sample text element is located as a piece of training data in the negative samples, and use the sample text element corresponding to the label 1 and the image corresponding to the row where the sample text element is located as a piece of training data in the positive samples, realizing the batch determination of multiple pieces of positive and negative sample training data.

[0278] For example, as Figure 11D shown, the text intersection matrix 1 and the sample matrix are intersected to obtain text intersection matrix 2. After that, the binary classification module can determine whether the intersection elements in text intersection matrix 2 are empty, thereby determining the positive and negative sample label matrix corresponding to the sample matrix. As Figure 11DThe label in the first row and first column of the positive and negative sample label matrix is 1. The element in the sample matrix corresponding to this label is the sample text element "1" in the first row and first column. This "1" and image a form a positive sample. As Figure 11D The label in the second row and first column of the positive and negative sample label matrix is 0. The sample text element in the second row and first column of the sample matrix corresponding to this label is "1". This "1" and image b form a negative sample.

[0279] It should be noted that generally, the elements on the diagonal of the sample matrix and the sample images corresponding to their rows form positive samples.

[0280] In the embodiments of the present application, the binary classification module obtains the text intersection matrix 2 by intersecting the text intersection matrix 1 and the sample matrix, so as to quickly determine positive and negative samples by using whether the elements in the text intersection matrix 2 are empty, improve the determination efficiency of positive and negative samples, and ensure the comprehensiveness and integrity of the distinction between positive and negative samples, enabling the model to be trained quickly.

[0281] The process of determining positive and negative samples is introduced above. Next, the process of training a binary classification model using positive and negative samples will be continued.

[0282] S9. The binary classification module trains the text-image matching model using positive and negative samples to obtain a trained text-image matching model.

[0283] Exemplarily, the text-image matching model can be a multilayer perceptron (MLP) model. As Figure 11E shown, the binary classification module processes the sample text elements in positive and negative samples using a text encoder to obtain text feature vectors of the sample text elements. And, the binary classification module processes the sample images in positive and negative samples using an image encoder to obtain image feature vectors of the sample images. Then, for each sample in positive and negative samples, the binary classification module can concatenate the text feature vector of the sample text element in the sample and the image feature vector of the sample image in the sample. Then, the binary classification module can input the concatenated text feature vector and image feature vector into the MLP model to train the MLP model. During the training process, the binary classification module can use a loss function to test the difference between the predicted value output by the trained MLP model and the actual value. The predicted value represents whether the predicted sample image and text match. The actual value represents the actual matching situation of the sample image and text.

[0284] In some embodiments, the above loss function can be a focal loss function. Specifically, the focal loss function can be FL(p t )=-at 1(1 - p t ) γ log(p t ). Wherein, pt represents the probability of the matching between the predicted sample image and its corresponding text by the image - text matching model, that is, the probability that the predicted sample image and its corresponding text are positive samples. at1 is a factor for adjusting the weights of positive and negative samples, which can be set according to the number of positive and negative samples to adjust the imbalance between the number of positive and negative samples. For example, at1 is 0.1. γ is a adjustment factor used to reduce the loss contribution of easily distinguishable samples. It should be understood that the larger pt is, the smaller the value of the loss function is.

[0285] Optionally, the above - mentioned image encoder can be a stacked auto - encoder, and the text encoder can be a count vectorizer.

[0286] It should be noted that the above - introduced focal loss function is only an example of the loss function, and this loss function can also be other types of loss functions. For example, this loss function is a loss function of the bce class. In addition, the above - mentioned MLP model is only an example of the image - text matching model, and this image - text matching model can also be other deep - learning models, and the present application does not limit it.

[0287] In some embodiments, the above - mentioned positive and negative sample label matrix can be used when calculating the value of the loss function.

[0288] In some embodiments, the use of the image - text matching model to perform secondary confirmation on candidate visual media is only an example. This image - text matching model can also be directly used for searching visual media. For example, after the user inputs a search statement, the image - text matching model can directly use the text feature vector of this search statement (or filtered search statement) and the image feature vectors of each visual media on the mobile phone to determine whether the visual media matches the search statement and obtain the corresponding boolean value.

[0289] The process of using the binary classification module to determine whether the candidate visual media matches the filtered search statement is introduced above. Next, the process of using the dynamic threshold module to determine whether the candidate visual media matches the filtered search statement will be continued in combination with S417c.

[0290] S417c: The binary classification module inputs the filtered search statement and the similarity between each candidate visual media and the filtered search statement into the dynamic threshold module, and obtains the boolean value 2 corresponding to each candidate visual media output by the dynamic threshold module.

[0291] In the embodiments of the present application, generally, the similarity threshold 1 corresponding to search statements of different lengths is different. The longer the length of the search statement, the higher the similarity degree between the visual content of the visual media and the search statement is required to be, and correspondingly, the similarity threshold 1 needs to be higher. Therefore, the secondary confirmation module can use the dynamic threshold module to determine a dynamic threshold that matches the length of the filtered search statement, that is, to determine the similarity threshold 1 corresponding to the filtered search statement, so as to determine the Boolean value 2 corresponding to each candidate visual media by using the similarity threshold 1 corresponding to the filtered search statement.

[0292] Exemplarily, as Figure 12 shown, the process by which the dynamic threshold module determines the Boolean value 2 corresponding to each candidate visual media may include: First, the dynamic threshold module can determine the length of the filtered search statement through the filtered search statement. After that, based on the length of the filtered search statement, the dynamic threshold model combines t = parameter 1 * L + parameter 2 to determine the similarity threshold 1 corresponding to the filtered search statement. Among them, the above parameter 1 and parameter 2 are preset parameters. For example, parameter 1 is 0.05 and parameter 2 is 0.33. Optionally, the dynamic threshold module can also determine the length of the filtered search statement through the word segmentation result of the filtered search statement.

[0293] After that, for each of the m candidate visual media, the dynamic threshold module can compare the size between the similarity between the candidate visual media and the filtered search statement and the similarity threshold 1 corresponding to the filtered search statement. In the case where the similarity between the candidate visual media and the filtered search statement is less than the similarity threshold 1, it indicates that the similarity degree between the visual content of the candidate visual media and the filtered search statement is relatively low, and the dynamic threshold module can determine that the Boolean value 2 corresponding to the candidate visual media is false.

[0294] In the case where the similarity between the candidate visual media and the filtered search statement is greater than or equal to the similarity threshold 1, it indicates that the similarity degree between the visual content of the candidate visual media and the filtered search statement is relatively high, and the dynamic threshold module can determine that the Boolean value 2 corresponding to the candidate visual media is true.

[0295] Among them, optionally, as shown above Figure 12 shown, the value range of t can be greater than or equal to parameter 3, that is, min(t, parameter 3), and the value range of t can be less than or equal to parameter 4. That is, max(t, parameter 4). Among them, parameter 3 and parameter 4 are preset.

[0296] In some embodiments, since the dynamic threshold module corresponding to S417c belongs to the branch corresponding to the base model, the similarity between the candidate visual media in S417c and the filtered search statement includes the similarity between the candidate visual media corresponding to the base model and the filtered search statement (i.e., similarity 1). Accordingly, the dynamic threshold module can determine whether the similarity 1 between the candidate visual media and the filtered search statement is less than the similarity threshold 1. When the similarity 1 is greater than or equal to the similarity threshold 1, the dynamic threshold module can determine that the Boolean value 2 corresponding to the candidate visual media is true. In the case where the similarity 1 is less than the similarity threshold 1, the dynamic threshold module can determine that the Boolean value 2 corresponding to the candidate visual media is false.

[0297] It can be understood that if the input parameter of the secondary confirmation module only includes one similarity between the candidate visual media and the filtered search statement, and does not include the similarity between the candidate visual media corresponding to the fine-tuning model and the filtered search statement and the similarity between the candidate visual media corresponding to the base model and the filtered search statement, then the image feature vector of the candidate visual media in S417c is the similarity between the input candidate visual media and the filtered search statement.

[0298] In addition, the similarity between the candidate visual media and the filtered search statement can also be determined by other multi-modal models, and the present application does not limit it.

[0299] The process of determining the Boolean value corresponding to the candidate visual media using the dynamic threshold module is introduced above. Next, the process of determining whether the candidate visual media matches the filtered search statement using the label confirmation module will be continued with reference to S417d.

[0300] S417d. The secondary confirmation module inputs the tokenization result of the above filtered search statement, label 1 included in the filtered search statement, and label 2 of each candidate visual media into the label confirmation module, and obtains the Boolean value 3 corresponding to each candidate visual media output by the label confirmation module.

[0301] In the embodiments of the present application, the secondary confirmation module can use the label confirmation module to determine whether label 2 of the candidate visual media exists in the tokenization of the filtered search statement or label 1 included in the filtered search statement, so as to determine whether the candidate visual media matches the filtered search statement, and thus determine the Boolean value 3 corresponding to the candidate visual media. Exemplarily, as Figure 13 shown, the process of the label confirmation model determining the Boolean value 3 corresponding to the candidate visual media can include:

[0302] First, for each of the m candidate visual media, the label confirmation module can determine whether the label 2 of the candidate visual media includes the word segments in the word segmentation result of the filtered search statement, and determine whether the label 2 of the candidate visual media includes the label 1 included in the filtered search statement, that is, determine whether the candidate visual media includes the visual content corresponding to the filtered search statement.

[0303] In the case where the label 2 of the above candidate visual media includes at least one word segment of the filtered search statement, or the label 2 of the candidate visual media includes at least one label 1 included in the filtered search statement, it indicates that the candidate visual media hits the filtered search statement. The label confirmation module can determine that the candidate visual media may be the visual media required by the user. Therefore, the label confirmation module can determine the Boolean value 3 corresponding to the candidate visual media as true. For example, the label 1 of the filtered search statement includes "child" and "flower", and the word segmentation result of the filtered search statement includes the word segments "child", "holding", and "flower". Then, in the case where the label 2 of the candidate visual media includes "child", "holding" or "flower", or the label 2 of the candidate visual media includes "child" or "flower", it is determined that the Boolean value 3 corresponding to the candidate visual media is determined as true.

[0304] In the case where the label 2 of the above candidate visual media does not include all the word segments of the filtered search statement, and the label 2 of the candidate visual media does not include all the label 1 included in the filtered search statement, it indicates that the candidate visual media does not hit the filtered search statement, and the candidate visual media may not be the visual media required by the user. Therefore, the label confirmation module can determine the Boolean value 3 corresponding to the candidate visual media as false, so as to obtain the Boolean value 3 of the m candidate visual media.

[0305] Optionally, in order to improve the accuracy of the search results, the label confirmation module can determine that the Boolean value 3 corresponding to the candidate visual media is true when the label 2 of the candidate visual media includes all the word segments of the filtered search statement, or includes all the label 1 included in the filtered search statement.

[0306] In the case where the label 2 of the candidate visual media does not include at least one word segment of the filtered search statement and does not include at least one label 1 included in the filtered search statement, it is determined that the Boolean value 3 corresponding to the candidate visual media is false.

[0307] The process of the secondary confirmation module using the label confirmation module to determine the Boolean value 3 corresponding to each candidate visual media is introduced above. Next, the process of the secondary determination module using the whitelist threshold module to determine whether the candidate visual media matches the filtered search statement will be continued.

[0308] The secondary confirmation module inputs the similarity between the filtered search statement and the candidate visual media into the whitelist threshold module, and obtains the boolean value 4 corresponding to each candidate visual media output by the whitelist threshold module.

[0309] In the embodiments of the present application, the secondary confirmation module can determine whether the filtered search statement hits the dictionary through the whitelist threshold module, so as to determine whether there is a similarity threshold 2 corresponding to the filtered search statement, and achieve the accurate determination of the similarity threshold. Exemplarily, the whitelist threshold module can determine whether the dictionary contains the filtered search statement. If it exists, the whitelist threshold module becomes effective, and the whitelist threshold module can use the similarity threshold corresponding to the filtered search statement in the dictionary as the similarity threshold 2 corresponding to the filtered search statement, so as to screen the candidate visual media that matches the filtered search statement by using the similarity threshold 2. If it does not exist, the whitelist threshold module does not become effective, that is, there is no need to use the whitelist threshold module to determine the boolean value 4 corresponding to each candidate visual media. Wherein, the dictionary includes a key and the value corresponding to the key, the key represents a preset search text, the value is the similarity threshold corresponding to the preset search text, and the key in the dictionary can be the search text frequently input by the user, that is, the search statement that has been pre-tested.

[0310] In some embodiments, there are corresponding dictionaries for the branches corresponding to the fine-tuning model and the base model. As Figure 14A shown, when the filtered search statement is the branch corresponding to the base model, the whitelist threshold module can determine whether the dictionary corresponding to the base model includes the filtered search statement, that is, determine whether there is a key in the dictionary corresponding to the base model that is the same as the filtered search statement.

[0311] If the dictionary corresponding to the base model includes the filtered search statement, it indicates that the filtered search statement hits the dictionary corresponding to the base model, that is, there is a key in the dictionary corresponding to the base model that is the same as the filtered search statement, then the whitelist threshold module can use the value corresponding to the filtered search statement in the dictionary corresponding to the base model as the similarity threshold 2.

[0312] After that, for each of the m candidate visual media, the whitelist threshold module can compare the similarity between the candidate visual media corresponding to the base model and the filtered search threshold with the similarity threshold 2. When the similarity between the candidate visual media corresponding to the base model and the filtered search threshold is less than the similarity threshold 2, the whitelist threshold module can determine that the boolean value 4 corresponding to the candidate visual media is false. When the similarity is greater than or equal to the similarity threshold 2, it indicates that the visual content of the candidate visual media is highly similar to the filtered search statement, and the whitelist threshold module can determine that the boolean value corresponding to the candidate visual media is true.

[0313] Correspondingly, when filtering the search statement for the branch corresponding to the base model, the above S418 can be that for each candidate visual media, when the boolean value corresponding to the candidate visual media has a true value, the secondary confirmation module (such as the fusion module in the secondary confirmation module) can use the candidate visual media as Visual Media 1. Among them, the boolean value corresponding to the candidate visual media includes the above-mentioned boolean value 1, boolean value 2, boolean value 3, and boolean value 4. That is to say, the secondary confirmation module can take the union of the boolean value 1, boolean value 2, boolean value 3, and boolean value 4 corresponding to the candidate visual media to obtain the boolean value corresponding to the candidate visual media.

[0314] When the boolean value corresponding to the candidate visual media is all false, it indicates that the boolean value 1, boolean value 2, boolean value 3, and boolean value 4 corresponding to the candidate visual media are all false. The secondary confirmation module (such as the fusion module in the secondary confirmation module) can not use the candidate visual media as Visual Media 1.

[0315] It should be noted that since both the above whitelist threshold module and the above dynamic threshold module are used to confirm the similarity threshold corresponding to the filtered search statement, and the similarity threshold determined by the whitelist threshold module is more accurate. Therefore, when the above whitelist threshold module takes effect, that is, when using the similarity threshold 2 corresponding to the filtered search statement to determine the boolean value corresponding to the candidate visual media, the dynamic threshold module can be ineffective, and the secondary confirmation module can not use the similarity threshold 1 corresponding to the filtered search statement to determine the boolean value corresponding to the candidate visual media, avoiding unnecessary determination of the similarity threshold, and thus avoiding unnecessary screening of candidate visual media.

[0316] In some embodiments, the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module in the above secondary confirmation module can determine the boolean value corresponding to the candidate visual media in parallel or serially. Whether it is determined in parallel or serially, when one module in the secondary determination module determines that the boolean value corresponding to the candidate visual media is true, other modules do not need to continue to determine the boolean value corresponding to the candidate visual media, that is, there is no need to input the relevant information of the candidate visual media into other models, avoiding unnecessary determination of the boolean value and ensuring the transmission cost.

[0317] In addition, the above secondary confirmation module including the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module is only an example. The secondary confirmation module may include one or more of the binary classification module, the dynamic threshold module, the label confirmation model, and the whitelist threshold module to improve the search efficiency of visual media. For example, if the secondary confirmation module includes one of the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module, such as the secondary confirmation module includes the binary classification module, correspondingly, the Boolean value corresponding to the above candidate visual media may include Boolean value 1. Another example is that the secondary confirmation module includes two of the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module, such as the secondary confirmation module includes the binary classification module and the dynamic threshold module. Correspondingly, the Boolean values corresponding to the above candidate visual media may include Boolean value 1 and Boolean value 2. Another example is that the secondary confirmation module includes three of the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module, such as the secondary confirmation module includes the binary classification module, the dynamic threshold module, and the label confirmation module. Correspondingly, the Boolean values corresponding to the above candidate visual media may include Boolean value 1, Boolean value 2, and Boolean value 3.

[0318] Moreover, the content included in the image information of the above candidate visual media and the information of the filtered search statement, that is, the input parameters of the secondary confirmation module, is only an example and can be adaptively set according to the modules included in the secondary confirmation module. For example, if the secondary confirmation module includes a binary classification model, the image information of the above candidate visual media may include the image feature vector of the candidate visual media, and the information of the filtered search statement may include the text feature vector of the filtered search statement.

[0319] The process in which the secondary confirmation module can sequentially use the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module to determine the Boolean value corresponding to the candidate visual media in the case where the filtered search statement corresponds to the branch of the base model is introduced above. Next, the process in which the secondary confirmation module can sequentially use the label confirmation module and the whitelist threshold module to determine the Boolean value corresponding to the candidate visual media in the case where the filtered search statement corresponds to the branch of the fine-tuning model will be continued.

[0320] S417f. The secondary confirmation module inputs the word segmentation result of the above filtered search statement, label 1 included in the filtered search statement, and label 2 of each candidate visual media into the label confirmation module to obtain the Boolean value 5 corresponding to each candidate visual media output by the label confirmation module.

[0321] Among them, the implementation process of S417f can refer to the implementation process of the above S417d and will not be elaborated here.

[0322] S417g. The secondary confirmation module inputs the similarity between the filtered search statement and the candidate visual media into the whitelist threshold module, and obtains the Boolean value 4 corresponding to each candidate visual media output by the whitelist threshold module.

[0323] Among them, the implementation process of S417g can refer to the implementation process of the above S417e. As Figure 14B shown, when the filtered search statement is the branch corresponding to the fine-tuning model, the whitelist threshold module can determine whether the dictionary corresponding to the fine-tuning model includes the filtered search statement, that is, determine whether there is a key in the dictionary corresponding to the fine-tuning model that is the same as the filtered search statement, so as to determine whether the whitelist threshold module takes effect.

[0324] After the whitelist threshold module takes effect, the secondary confirmation module can compare the similarity between the candidate visual media corresponding to the fine-tuning model and the filtered search statement with the value corresponding to the filtered search statement in the dictionary corresponding to the fine-tuning model to determine the Boolean value corresponding to the candidate visual media.

[0325] Correspondingly, when the filtered search statement is the branch corresponding to the fine-tuning model, the above S418 can be that for each candidate visual media, when the Boolean value corresponding to the candidate visual media is true, the secondary confirmation module (such as the fusion module in the secondary confirmation module) can use the candidate visual media as Visual Media 1. Among them, the Boolean value corresponding to the candidate visual media includes the above Boolean value 5 and the above Boolean value 6, that is to say, the secondary confirmation module can take the union of the Boolean value 5 and the Boolean value 6 corresponding to the candidate visual media to obtain the Boolean value corresponding to the candidate visual media.

[0326] In the embodiments of the present application, after obtaining the candidate visual media, the mobile phone uses the secondary confirmation module to continue to determine the Boolean value corresponding to the candidate visual media to determine whether the candidate visual media matches the filtered search statement, so as to select the candidate visual media that matches the filtered search statement from the candidate visual media, obtain the corresponding search result, ensure the accuracy of the search result, and thus ensure the user satisfaction.

[0327] It should be noted that the operations performed by the above modules or models are only examples, and the operations performed by the above modules can also be performed by other modules in the mobile phone, and the present application does not limit it. In addition, the operations actually performed by the above modules or models are performed by the mobile phone.

[0328] In some embodiments, when the mobile phone determines the candidate visual media or performs secondary confirmation on the candidate visual media, it may not filter the non-visual semantic entities in the search statement first, but directly use the text feature vector of the search statement to determine the candidate visual media, or perform secondary confirmation on the candidate media.

[0329] In some embodiments, the search and storage of the above visual media are carried out under the authorization of the user, including but not limited to notifying and reminding the user to read the relevant user agreement (notification) before the user uses this function, and signing the agreement (authorization) including authorizing the relevant user information.

[0330] In some embodiments, the present application provides a computer storage medium including computer instructions, which when running on an electronic device, cause the electronic device to execute the method as described above.

[0331] In some embodiments, the present application provides a computer program product, which when running on an electronic device, causes the electronic device to execute the method as described above.

[0332] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part.

[0333] It should be understood that the "embodiments" mentioned throughout the specification mean that the specific features, structures or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, the various embodiments mentioned throughout the specification do not necessarily refer to the same embodiments. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0334] It should also be understood that in the present application, "when...", "if" and "in case" all mean that under certain objective circumstances, the UE or the base station will make corresponding processing, which does not limit the time, and does not require the UE or the base station to have a judgment action when implemented, nor does it mean that there are other limitations.

[0335] Those of ordinary skill in the art can understand that the various numerical numbers such as the first and second involved in this application are only for the convenience of description and are not used to limit the scope of the embodiments of this application, nor do they represent the order of precedence. In this application, the elements represented in the singular are intended to mean "one or more", rather than "one and only one", unless otherwise specified. In this application, unless otherwise specified, "at least one" is intended to mean "one or more", and "a plurality" is intended to mean "two or more".

[0336] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A can be singular or plural, and B can be singular or plural. The term "at least one of..." or "at least one kind of..." in this article means all or any combination of the items listed. For example, "at least one of A, B, and C" can represent: A exists alone, B exists alone, C exists alone, A and B exist simultaneously, B and C exist simultaneously, and A, B, and C exist simultaneously. Among them, A can be singular or plural, B can be singular or plural, and C can be singular or plural.

[0337] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0338] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0339] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0340] In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, may exist independently physically for each unit, or two or more units may be integrated in one unit.

[0341] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0342] The same or similar parts among the various embodiments in the present application can be referred to each other. In the various embodiments in the present application, and in each implementation manner / implementation method / realization method in each embodiment, if there is no special description and logical conflict, the terms and / or descriptions between different embodiments, and between each implementation manner / implementation method / realization method in each embodiment, are consistent and can be referenced to each other. The technical features in different embodiments, and in each implementation manner / implementation method / realization method in each embodiment, can be combined according to their internal logical relationships to form new embodiments, implementation manners, implementation methods, or realization methods. The implementation manners of the present application described above do not constitute a limitation on the protection scope of the present application.

[0343] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claimed rights. In short, the above description is only a preferred embodiment of the technical solution of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A visual media search method applicable to an electronic device, characterized in that, the method includes: displaying a first interface; the first interface includes a search box; receiving a search statement input in the search box; displaying search results, the search results corresponding to a first visual media; the first visual media represents a visual media whose visual content matches the search statement; wherein, the first visual media is a visual media whose first similarity with the search statement is greater than a first threshold, or the first visual media is a visual media whose second similarity with the search statement is greater than a second threshold; the first similarity is the similarity between the visual media and the search statement determined by a base model, and the second similarity is the similarity between the visual media and the search statement determined by a fine-tuning model, and the fine-tuning model is obtained by training the base model based on a training set corresponding to a preset search scenario.

2. The method according to claim 1, characterized in that, when the search statement does not belong to a first preset whitelist, the first visual media is a visual media whose first similarity with the search statement is greater than the first threshold; when the search statement belongs to the first preset whitelist, the first visual media is a visual media whose second similarity with the search statement is greater than the second threshold.

3. The method according to claim 1 or 2, characterized in that, the method further includes: when each semantic entity in the search statement belongs to the first preset whitelist, determining that the search statement belongs to the first preset whitelist; when there is a semantic entity in the search statement that does not belong to the first preset whitelist, determining that the search statement does not belong to the first preset whitelist.

4. The method according to any one of claims 1 to 3, characterized in that, the method further includes: obtaining a first image feature vector of the visual media determined by the base model and a second image feature vector of the visual media determined by the fine-tuning model from the electronic device; obtaining the first similarity between the visual media and the search statement based on the first image feature vector and a first text feature vector of the search statement determined by the base model; obtaining the second similarity between the visual media and the search statement based on the second image feature vector and a second text feature vector of the search statement determined by the fine-tuning model.

5. The method according to claim 4, characterized in that, the base model and the fine-tuning model share an image encoder, the first image feature vector includes a first feature vector of the visual media output by the image encoder and a first L2 norm of the visual media, and the first L2 norm of the visual media is determined based on the first feature vector of the visual media and a first mapping matrix in the base model; the first mapping matrix is used to reduce the dimension of the first feature vector; The second image feature vector includes the first feature vector and the second L2 norm of the visual media, where the second L2 norm of the visual media is determined based on the first feature vector of the visual media and the second mapping matrix in the fine-tuning model; the second mapping matrix is used to reduce the dimension of the first feature vector.

6. The method according to claim 4 or 5, wherein, the base model and the fine-tuning model share a text encoder; the first text feature vector is determined based on the second feature vector of the search statement output by the text encoder and the third mapping matrix in the base model; the second text feature vector is determined based on the second feature vector and the fourth mapping matrix in the fine-tuning model.

7. The method according to claim 5, wherein, the method further includes: storing the first feature vector, the first L2 norm, and the second L2 norm of the visual media in the electronic device.

8. The method according to any one of claims 5 to 7, wherein, The first similarity between the visual media and the search statement is determined using ; the S 1 represents the first similarity, α 1 represents the first L2 norm of the visual media, T 1 ' represents the first text feature vector of the search statement, and the X represents the first feature vector of the visual media.

9. The method according to any one of claims 1 to 8, wherein, the first visual media is a visual media whose first similarity with the filtered search statement is greater than the first threshold, or the first visual media is a visual media whose second similarity with the filtered search statement is greater than the second threshold; the filtered search statement is obtained by filtering non-visual semantic entities in the search statement, and the non-visual semantic entities do not correspond to visual content.

10. An electronic device, wherein, the electronic device includes a display screen, a memory, and one or more processors; the display screen, the memory, and the processor are coupled; the display screen is used to display images generated by the processor, the memory is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the electronic device performs the method according to any one of claims 1 to 9.

11. A computer storage medium, wherein, including computer instructions, when the computer instructions run on an electronic device, the electronic device performs the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Searching method and system

    CN102456054A

  • Medical title matching method, device, equipment and storage medium

    CN113569124A

  • Visual media personalized search method and device

    CN113641857A

  • Blacklist fuzzy matching method based on character string text feature similarity

    CN115577269A

  • Text generation method, place retrieval method and related devices

    CN116662583A