Visual media searching method and electronic equipment

By generating the sample matrix and training the graphic matching model, the problem of inaccurate search of complex search statements in the prior art is solved, and high-accuracy search of visual media is achieved.

CN120067395AActive Publication Date: 2025-05-30HONOR DEVICE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202311589198.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-30
Estimated Expiration
2043-11-23

AI Technical Summary

Technical Problem

In the prior art, when processing complex search statements, it is difficult to accurately retrieve visual media, resulting in inaccurate search results.

Method used

By generating a sample matrix, matching the sample images with their corresponding text elements, and training the graphic matching model to improve the accuracy of the search results.

Benefits of technology

It realizes that no matter how complex the search statement is, visual media can be accurately retrieved and meet users' search needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067395A_ABST
    Figure CN120067395A_ABST
Patent Text Reader

Abstract

The invention provides a visual media searching method and electronic equipment, and relates to the technical field of image processing. After receiving a search statement input by a user, the electronic equipment determines a text feature vector of the search statement. Afterwards, the electronic equipment can input the text feature vector of the search statement and the image feature vector of the visual media locally stored in the electronic equipment into an image-text matching model, so that whether the search statement is matched with the visual media or not is determined through the image-text matching model, and search of the visual media is achieved. Wherein the image-text matching model is obtained by training positive and negative samples, the positive and negative samples are determined by distinguishing sample text elements corresponding to sample images, and the sample text elements corresponding to the sample images comprise one text element in a text set corresponding to each sample image; the text set corresponding to the sample image represents the set of the description content corresponding to the sample image, and rapid determination of the training sample of the image-text matching model is realized, so that the training efficiency of the image-text matching model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular, to a visual media search method and an electronic device. Background Art

[0002] With the development of electronic devices (such as mobile phones), the shooting function of mobile phones has also developed rapidly. More and more users use mobile phones to take photos, videos, etc., and store the taken photos and videos in the photo gallery of the mobile phone. In addition, users can also store the pictures and screenshots downloaded by the mobile phone in the photo gallery of the mobile phone.

[0003] When a user wants to search for visual media (such as photos, videos, etc.), the user can enter a search statement on the mobile phone. For example, for the photos taken on September 1st, the mobile phone responds to the search statement entered by the user, retrieves the photos taken on September 1st, and obtains corresponding search results. However, the mobile phone's ability to understand search statements is limited. When the search statement is relatively complex, the mobile phone may not be able to accurately obtain the corresponding search results. Summary of the Invention

[0004] In view of this, this application provides a visual media search method and an electronic device for improving the accuracy of search results.

[0005] In a first aspect, this application provides a visual media search method applied to an electronic device. The electronic device can obtain P sample data pairs, and the P sample data pairs include P sample images and the description texts corresponding to the P sample images.

[0006] After that, for each sample image among the P sample images, a text element is selected from the text set corresponding to the sample image to obtain the first text element (or referred to as text element 1) corresponding to the sample image. Among them, the text set corresponding to the sample image may include the description text corresponding to the sample image and each sub-text (or referred to as the first sub-text) in the description text. The first text element may be the description text corresponding to the sample image or a sub-text in the description text.

[0007] After that, the electronic device can perform image-text matching on the P sample images and the first text elements corresponding to the P sample images, that is, match any sample image with the first text element corresponding to any sample image to obtain a P*P sample matrix. Among them, the element in the i-th row of the sample matrix represents each sample text element corresponding to the i-th sample image; any sample text element in the sample matrix and the sample image corresponding to the row where the sample text element is located form a sample.

[0008] After that, the electronic device can train a graphic-text matching model based on the sample matrix and the sample images.

[0009] The electronic device can display a first interface; the first interface includes a search box;

[0010] Receive a search statement entered in the search box; the search statement includes one or more sub - texts;

[0011] Display search results, where the search results correspond to a first visual medium; the first visual medium represents a visual medium whose visual content determined by a graphic - text matching model matches the search statement.

[0012] In this application, the electronic device can obtain P pairs of sample data, and these P sample data are P positive samples. Then, the electronic device can use the P pairs of data to generate a P*P sample matrix, thereby obtaining P*P samples, realizing the increment of the sample quantity, realizing the rapid generation of training samples, and ensuring the training effect of the graphic - text matching model. Then, the electronic device can use the graphic - text matching model to search for a visual medium that matches the search statement input by the user, so that no matter how complex the search statement is, the search accuracy of the visual medium can be ensured to meet the user's search needs. In addition, in the scenario of graphic - text search, considering that the search statement input by the user conforms to the expression habits of natural language, there may be multiple pieces of information (i.e., the second sub - text), and this second sub - text is not a simple combination. Therefore, when training the graphic - text matching model, the sample image can also have content describing the sample image, forming a text set corresponding to the sample image. The text set corresponding to the sample image can include both descriptions of the overall visual content of the sample image and descriptions of partial visual content of the sample image, so as to train a graphic - text matching model using the text set corresponding to the sample image, so that no matter whether the search statement input by the user corresponds to the overall visual content of the visual medium or the partial visual content of the visual content, the matching between the visual medium and the search statement can be realized.

[0013] In a possible design, the process of determining the P*P sample matrix can include:

[0014] For each sample image, use the first text elements corresponding to the P sample images as the sample text elements corresponding to the sample image.

[0015] Generate a sample matrix based on the sample text elements corresponding to each sample image. Among them, the i - th row element in the sample matrix represents the respective sample text elements corresponding to the i - th sample image. Any sample text element in the sample matrix and the sample image corresponding to the row where the sample text element is located form a sample. Based on this, the matching of the P sample images and the P first text elements is realized, obtaining multiple samples, and ensuring the rapid generation of samples.

[0016] In a possible design, the process of training the text-image matching model based on the sample matrix and the sample image may include:

[0017] The electronic device can distinguish whether the sample text element in the sample matrix belongs to the positive sample or the negative sample according to the sample matrix and the text set corresponding to the sample image.

[0018] After that, the electronic device can use the sample text elements and the sample image in the sample matrix to obtain positive and negative samples. Then, the electronic device can train a text-image matching model based on the positive and negative samples to achieve the distinction between positive and negative samples.

[0019] In a possible design, the process of determining the above positive and negative samples may include:

[0020] For each sample text element in the sample matrix, the electronic device can determine whether the sample text element belongs to the text set corresponding to the sample image corresponding to the row where the sample text element is located, so as to determine whether the sample text element matches the sample image.

[0021] When the sample text element belongs to the text set corresponding to the sample image corresponding to the row where the sample text element is located, it indicates that the sample element matches the sample image, and the electronic device determines that the sample text element and the sample image corresponding to the row where the sample text element is located belong to the positive sample;

[0022] When the sample text element does not belong to the text set corresponding to the sample image corresponding to the row where the sample text element is located, it indicates that the sample element matches the sample image, and the electronic device can determine that the sample text element and the sample image corresponding to the row where the sample text element is located belong to the negative sample, realizing the distinction between positive and negative samples.

[0023] In another possible design, the process of determining the above positive and negative samples may include:

[0024] Find the intersection of the first text intersection matrix (or text intersection matrix 1) and the sample matrix to obtain the second text intersection matrix (or text intersection matrix 2); the first text intersection matrix is determined based on the same elements in the text sets corresponding to any two sample images.

[0025] When the s-th intersection element in the t-th row of the second text intersection matrix is empty, it is determined that the s-th intersection element in the t-th row of the sample matrix and the t-th sample image belong to the negative sample.

[0026] When the s-th intersection element in the t-th row of the second text intersection matrix is not empty, it is determined that the s-th intersection element in the t-th row of the sample matrix and the t-th sample image belong to the positive samples. Based on this, by taking the intersection of the sample matrix and the first text intersection matrix, it is possible to quickly determine whether the text elements in the sample matrix belong to the positive samples, improving the efficiency of determining positive and negative samples.

[0027] In a possible design, for each sample image, the descriptive text corresponding to the sample image is tokenized to obtain the first sub-text in the descriptive text; wherein, the first sub-text includes nouns and / or phrases in the descriptive text. The phrase can be a descriptive phrase, thereby realizing the determination of the set of descriptive contents corresponding to the sample image.

[0028] In a possible design, after obtaining the text sets corresponding to the sample images, the electronic device determines the union of the text sets corresponding to the respective sample images. Thereafter, the electronic device assigns numbers to each text element in the union of the text sets such that different text elements in the text set correspond to different numbers and the same text element corresponds to the same number.

[0029] Correspondingly, the electronic device can perform graphic-text matching on the P sample images and the numbers corresponding to the first text elements corresponding to the P sample images to obtain a P*P sample matrix, improving the generation efficiency of the sample matrix.

[0030] In a possible design, the electronic device can randomly select text elements from the text set corresponding to each sample image according to a preset ratio; wherein, the preset ratio includes the ratio of the first text elements selected that are descriptive texts and the ratio of the first text elements selected that are first sub-texts.

[0031] In the present application in China, since generally the image encoder and the text encoder are trained using the image and its corresponding text as a whole, therefore, the ratio of the text elements selected in the preset ratio that are descriptive texts will be greater than the ratio of the text elements selected that are sub-texts, ensuring the training effect of the model.

[0032] In a possible design, the first visual media refers to the candidate visual media determined by the graphic-text matching model, where the visual content matches the search statement. The candidate visual media refers to the visual media whose similarity to the search statement is greater than the first threshold.

[0033] In the present application, the graphic-text matching model can be used for the secondary confirmation of the candidate visual media, that is, to further determine whether the candidate visual media preliminarily retrieved by the electronic device matches the search statement.

[0034] In a possible design, after obtaining the search statement, the electronic device can filter the search statement to filter out the non-visual semantic entities in the search statement, obtaining a filtered search statement. Subsequently, the electronic device can use an image-text matching model to determine a matching value (such as a boolean value) corresponding to the visual media based on the image information of the visual media and the information of the filtered search statement.

[0035] Determine the first visual media according to the matching value corresponding to the visual media. Moreover, the electronic device can determine a second visual media (or referred to as visual media 2) that matches the non-visual semantic entity in the search statement.

[0036] Subsequently, the electronic device can use the intersection between the first visual media and the second visual media as the target visual media to accurately determine the search result.

[0037] Among them, the matching value corresponding to the above visual media can include a boolean value (or referred to as boolean value 1). When the boolean value corresponding to the visual media is true, it is determined that the visual media is the first visual media.

[0038] When the boolean value corresponding to the visual media is false, it is determined that the visual media is not the first visual media.

[0039] In a possible design, in response to an operation of opening the gallery application, display the first interface;

[0040] Or, in response to an operation of opening the negative first screen triggered on the home screen of the electronic device, display the first interface;

[0041] Or, in response to a pull-down search operation triggered on the home screen of the electronic device, display the first interface.

[0042] In a possible design, the electronic device can input the image feature vectors of each of the visual media in the electronic device and the text feature vector of the search statement into the image-text matching model to obtain the matching results corresponding to each of the visual media. The image-text matching model is used to obtain the matching degree between the visual media and the search statement based on the image feature vector of the visual media and the text feature vector of the search statement, and compare the matching degree between the visual media and the search statement with a preset classification threshold to obtain the matching result corresponding to the visual media.

[0043] In a second aspect, the present application provides an electronic device, which includes a display screen, a memory, and one or more processors; the display screen, the memory, and the processors are coupled; the display screen is configured to display an image generated by the processors, the memory is configured to store computer program code, and the computer program code includes computer instructions; when the processors execute the computer instructions, the electronic device is caused to execute the method as described above.

[0044] In a third aspect, the present application provides a computer storage medium, including computer instructions, which when running on an electronic device, cause the electronic device to execute the method as described above.

[0045] In a fourth aspect, the present application provides a computer program product, which when running on an electronic device, causes the electronic device to execute the method as described above.

[0046] It can be understood that for the beneficial effects that can be achieved by the electronic device described in the second aspect, the computer storage medium described in the third aspect, and the computer program product described in the fourth aspect provided above, reference can be made to the beneficial effects in the first aspect and any of its possible design manners, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1A FIG. 1 is a schematic diagram of an interface for visual media search provided by an embodiment of the present application;

[0048] Figure 1B FIG. 2 is a schematic diagram of an interface for visual media search provided by an embodiment of the present application;

[0049] Figure 1C FIG. 3 is a schematic diagram of an interface for visual media search provided by an embodiment of the present application;

[0050] Figure 1D FIG. 4 is a schematic diagram of an interface for visual media search provided by an embodiment of the present application; Figure Four ;

[0051] Figure 2A FIG. 5 is a block diagram of the structure of an electronic device provided by an embodiment of the present application;

[0052] Figure 2B FIG. 6 is a software structure diagram of an electronic device provided by an embodiment of the present application;

[0053] Figure 3A FIG. 7 is a schematic diagram of an interface for visual media search provided by an embodiment of the present application;

[0054] Figure 3B FIG. 8 is a schematic diagram of an interface for visual media search provided by an embodiment of the present application;Figure Six ;

[0055] Figure 4 Schematic diagram one of a visual media search method provided by an embodiment of the present application;

[0056] Figure 5A Schematic diagram one of the determination process of an image feature vector provided by an embodiment of the present application;

[0057] Figure 5B Schematic diagram two of the determination process of an image feature vector provided by an embodiment of the present application;

[0058] Figure 5C Schematic diagram of the determination process of a text feature vector provided by an embodiment of the present application;

[0059] Figure 5D Schematic diagram of the determination process of a similarity provided by an embodiment of the present application;

[0060] Figure 6 Schematic diagram two of a visual media search method provided by an embodiment of the present application;

[0061] Figure 7 Schematic diagram three of a visual media search method provided by an embodiment of the present application;

[0062] Figure 8 Schematic diagram of a visual media search method provided by an embodiment of the present application Figure Four ;

[0063] Figure 9A Schematic diagram one of a secondary confirmation provided by an embodiment of the present application;

[0064] Figure 9B Schematic diagram one of a visual media provided by an embodiment of the present application;

[0065] Figure 10 Schematic diagram five of a visual media search method provided by an embodiment of the present application;

[0066] Figure 11A Schematic diagram one of a matrix provided by an embodiment of the present application;

[0067] Figure 11B Schematic diagram two of a visual media provided by an embodiment of the present application;

[0068] Figure 11C Schematic diagram two of a matrix provided by an embodiment of the present application;

[0069] Figure 11D Schematic diagram of the positive and negative sample discrimination process provided by an embodiment of the present application;

[0070] Figure 11E Schematic diagram of the training process of a binary classification model provided by an embodiment of the present application;

[0071] Figure 12 Schematic diagram two of a secondary confirmation provided by an embodiment of the present application;

[0072] Figure 13 Schematic diagram three of a secondary confirmation provided by an embodiment of the present application;

[0073] Figure 14A Schematic diagram of a secondary confirmation provided by an embodiment of the present application Figure Four ;

[0074] Figure 14B Schematic diagram five of a secondary confirmation provided by an embodiment of the present application. Specific implementation manners

[0075] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this embodiment, unless otherwise stated, the meaning of "a plurality" is two or more.

[0076] To understand the embodiments of the present application more clearly, the following first explains the vocabulary involved in the present application.

[0077] Visual media: refers to pictures or videos.

[0078] Semantic entity: The named entity recognition technology (NER) can identify statements and recognize entities with specific meanings in the text, such as personal names and place names. In this solution, the entities with specific meanings identified are called semantic entities.

[0079] Visual content-related and visual content-unrelated: Visual content refers to the objects presented by visual media and their interrelationships, etc. Simply put, visual content can be understood as the content included in visual media. In the context of image search in the present application, the data that can be obtained after the natural picture understanding of visual media files by the model is called "visual content-related". This solution refers to the data that is related to visual media files and can be obtained without the picture understanding ability of the model as "visual content-unrelated". For example, when an electronic device collects visual media files, it can obtain and save the shooting location, shooting time, name, file attributes, etc.

[0080] For example, in the phrase "photos taken in City 1 this year", "this year" (the time of shooting), "City 1" (the location of shooting), and "photos" (the file attribute) are all data that can be obtained and saved when the electronic device collects visual media files. Therefore, "this year", "City 1", and "photos" are not related to visual semantics. In the phrase "the sky photographed in City 1 this year", "the sky" can only be obtained through the image understanding ability of the model to understand the picture. Therefore, "the sky" is related to visual semantics.

[0081] Text semantic vector: It can be obtained by feeding the text into a text encoder, which is a vector that can represent the semantic features of the entire sentence. The text encoder can use the clip model or other models, such as the Transformer commonly used in natural language processing (NLP). This solution is not limited here. Among them, in this application, the text semantic vector can also be called the text feature vector.

[0082] Visual semantic vector: It can be obtained by feeding visual media (such as images) into an image encoder. The image encoder can use the clip model or other models, such as the CNN model or the VIT model. This solution is not limited here. Among them, in this application, the visual semantic vector can also be called the image feature vector.

[0083] Vector similarity: It is used to describe the similarity between two vectors (for example, between the text semantic vector and the visual semantic vector). In the embodiments of this application, the similarity between the text semantic vector of the search statement and the visual semantic vector of the visual media can be compared to determine the visual media that matches the search statement. Generally, the vector similarity can be calculated through the cosine similarity calculation formula. Of course, it can also be calculated through other methods.

[0084] An electronic device (such as a mobile phone) can manage the user's pictures, videos and other visual media through a gallery application. Taking the example of a mobile phone taking a photo, after the mobile phone takes a photo, the gallery application determines the attribute tags such as the shooting location, shooting time, and photo name corresponding to the photo and can use this attribute tag as the index of the picture. After the gallery application establishes an index for the visual media, it can provide corresponding search services to the user. Specifically, the user can enter keywords in the gallery application to search for pictures or videos on the mobile phone. Exemplarily, the user can enter keywords such as "sky", "cat", "time point 1" in the search box provided by the gallery application. The gallery application matches the keywords entered by the user with the indexes of the pictures, videos and other visual media in the gallery application to obtain the search results.

[0085] Optionally, the above property tags may further include attributes such as the face identity number (industrialdesign, ID) of the person in the photo, the person's name, and the relationship between the person and the mobile phone user. Among them, the face ID of the person in the photo can be automatically generated by the gallery application, and the name of the person in the photo and the relationship between the person and the mobile phone user can be manually input by the user. In practical applications, the same person corresponds to the same face ID, the same name, and the same relationship with the mobile phone user. Therefore, in order to simplify the user operation, the user only needs to input the person's name and the relationship with himself once for the same person. Subsequently, the gallery application automatically configures the person's face ID, name, and the relationship between the person and the mobile phone user for the pictures containing the person's face through face recognition technology. In addition, the above property information such as the photo name can also be automatically generated by the gallery application or manually named by the user.

[0086] The following exemplarily describes the interface involved in the search process of the gallery application with reference to the accompanying drawings:

[0087] As Figure 1A shown in (a) of Figure 1A , the mobile phone can display the main interface 101, which can also be called the desktop. The main interface 101 may include the icon 102 of the gallery application. The mobile phone receives the operation of the user clicking the icon 102. In response to this operation, the mobile phone can start the gallery application and display the interface 103 as shown in (b) of

[0088] As Figure 1A shown in (b) of

[0089] As Figure 1A shown in (b) of Figure 1AThe interface 105 shown in (c) therein can be referred to as a search interface. Among them, the interface 105 can display classification information of photos to the user. For example, in the interface 105, the mobile phone classifies the photos on the device according to time, portraits, and things, etc. For example, in the dimension of time, the mobile phone classifies the photos on the device into three time periods: "this month", "last month", and "this year". Among them, the "this month" album includes the photos or videos taken by the mobile phone this month, the "last month" album includes the photos or videos taken by the mobile phone last month, and the "this year" album includes the photos or videos taken by the mobile phone this year. In the dimension of portraits, the mobile phone classifies the photos on the device according to different people, such as the four different people in the interface 105. In the dimension of things, the mobile phone classifies and displays the photos on the device according to "scenery", "animals", "documents", and "buildings". It should be noted that the above classification dimensions can also be others, and no specific restrictions are made here. In the interface 105, the user can see this classification information without entering keywords.

[0090] Optionally, the interface 105 can also include options for search history 107 and "clear" 108. The search history includes the keywords that the user has entered, such as "flowers", "coffee", "cat", etc. The mobile phone can receive the operation of the user clicking "clear" 108, and in response to this operation, the mobile phone can clear the search history. After the mobile phone clears the search history, the keywords that the user has entered are no longer displayed on the search interface 105. For example, in response to the operation of the user clicking "clear" 108, as Figure 1B shown, the options for search history 107 and "clear" 108 are no longer displayed on the search interface 105, and the content displayed below moves up.

[0091] In response to the operation of the user entering the keyword "sky" in the interface 105, the mobile phone displays as Figure 1CThe interface 109 shown in (a) therein. Among them, the mobile phone can search for data related to the keyword "sky" on the local machine. Specifically, the mobile phone can associate with the keyword "sky" to obtain associated words such as "sky" and photos containing the word "sky". Then, search according to each associated word to obtain the search results of each associated word. For example: 100 photos related to "sky" and 32 photos related to photos containing the word "sky". Among them, the 100 photos related to "sky" can be recalled because the classification labels of each of these 100 photos match "sky" or "sky" and its associated words; the 32 photos related to photos containing the word "sky" can be recalled because through optical character recognition (OCR) technology, it is recognized that these 32 photos contain characters such as "sky". The union of the search results of each of these multiple associated words can be used as the search result of the keyword "sky".

[0092] The interface 109 also displays some search results of the keyword "sky" and the "More" option 110 corresponding to the search result of the keyword "sky". The mobile phone receives the click operation of the user on the "More" option 110 and displays as Figure 1C the interface 111 shown in (b) therein. Among them, the interface 111 is used to display photos and videos in the search results of the keyword "sky". Optionally, the photos and videos can be classified and displayed according to time. In addition, the interface 111 also includes a return key 112 and a title 113. In response to the operation of the user on the return key 112, the mobile phone can redisplay the interface 109. The title 113 may include the keyword "sky".

[0093] That is to say, in the gallery application, when the user enters a simple search statement in the search box, such as a simple keyword, for example: sky, location 1, time 1, etc., corresponding search results can be obtained. However, because the mobile phone's ability to understand and associate with search statements is limited, if the user enters a more complex search statement in the search box, if the keywords in the search statement cannot match the attribute labels of the pictures or the text in the pictures, it may not be able to find any photos. That is to say, the existing gallery applications do not support the search function based on complex search statements. As Figure 1D shown, when the user enters a more complex search statement "warming oneself around the stove and making tea" in the search box of the interface 114, the mobile phone cannot understand the associated words of "warming oneself around the stove and making tea", and since the photos do not have labels that can match "warming oneself around the stove and making tea" or its associated words, the search result shows "no pictures".

[0094] In some embodiments, an electronic device may utilize a graphic-text matching model to calculate the matching degree between visual media on the electronic device and a search statement, so that the electronic device can use the visual media with a matching degree higher than a threshold as a search result. Subsequently, the electronic device displays the search result to present the visual media required by the user. However, before using the graphic-text matching model to search for visual media, it is necessary to train the graphic-text matching model using training samples to obtain a graphic-text matching model that can accurately identify whether visual media is relevant to a search statement. However, training samples are generally obtained manually, resulting in a long time required to obtain training samples and low training efficiency.

[0095] Therefore, in order to improve the acquisition efficiency of training samples for the graphic-text matching model and thus improve the training efficiency of the graphic-text matching model, this application provides a solution for determining training samples (i.e., positive and negative samples) of the graphic-text matching model. The electronic device performs word segmentation on the description text corresponding to the sample image in a preset sample data pair to obtain sub-texts in the description text corresponding to the sample image. Subsequently, for each sample image, the electronic device can form a text set corresponding to the sample image by using the description text corresponding to the sample image and the sub-texts in the description text. Subsequently, for each sample image, the electronic device can randomly select a text element from the text set corresponding to the sample image as the text element 1 corresponding to the sample image. Subsequently, the electronic device can use the text element 1 corresponding to each sample image as the sample text element corresponding to the sample image, such that there are multiple sample text elements corresponding to one sample image, where one sample image and its corresponding one sample text element can form a sample, thereby obtaining multiple samples. Subsequently, the electronic device can distinguish positive and negative samples among the multiple samples, realizing the rapid determination of a large number of positive and negative samples and shortening the acquisition time of training samples for the graphic-text matching model.

[0096] Exemplarily, the above-mentioned electronic device can be a device capable of storing visual media such as a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc.

[0097] Exemplarily, Figure 2AThe structural schematic diagram of the electronic device 200 is shown. The electronic device 100 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone interface 270D, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.

[0098] Among them, the sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.

[0099] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0100] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0101] The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0102] A memory can also be provided in the processor 210 for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory can hold the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can directly call it from the said memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0103] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0104] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may also adopt different interface connection methods or combinations of multiple interface connection methods in the above embodiments.

[0105] The electronic device 200 realizes the display function through a GPU (Graphics Processing Unit), a display screen 294, and an application processor, etc. The GPU is connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs that execute program instructions to generate or change display information.

[0106] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 294, where N is a positive integer greater than 1.

[0107] The electronic device 200 can implement the shooting function through an ISP, a camera 293, a video codec, a GPU, a display screen 294, an application processor, etc.

[0108] The camera 293 is used to capture still images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB, YUV, etc. In some embodiments, the electronic device 100 may include one or N cameras 293, where N is a positive integer greater than 1.

[0109] The video codec is used to compress or decompress digital videos. The electronic device 200 can support one or more video codecs. In this way, the electronic device 200 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0110] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously learn on its own. Through the NPU, applications such as intelligent cognition of the electronic device 200 can be realized, such as: image recognition, face recognition, voice recognition, text understanding, etc.

[0111] The external memory interface 220 can be used to connect to an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 200. The external memory card communicates with the processor 210 through the external memory interface 220 to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.

[0112] The internal memory 221 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 221 can include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.). The data storage area can store the data created during the use of the electronic device 200 (such as audio data, phone book, etc.). In addition, the internal memory 221 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. The processor 210 executes various functional applications and data processing of the electronic device 200 by running the instructions stored in the internal memory 221 and / or the instructions stored in the memory provided in the processor.

[0113] Figure 2B It is the software structure block diagram of the electronic device 200 in the embodiments of the present application. The software system of the electronic device 200 can adopt a layered architecture. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate through software interfaces. Taking the Android system as an example, in some embodiments, the Android system can be divided into four layers, from top to bottom are the application layer, application framework layer, Android runtime, and system library, and the kernel layer.

[0114] As Figure 2B shown, the application layer can include application programs such as a gallery service module, a search module, a multimodal understanding module, a natural language understanding module, and a camera application.

[0115] The application framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer includes some predefined functions.

[0116] The system library can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (such as: OpenGL ES), 2D graphics engine (such as: SGL), etc. Among them, the media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0117] The kernel layer is the layer between hardware and software.

[0118] Next, in combination with the capture and photo-taking scenario, the working processes of the software and hardware of the electronic device 200 will be exemplarily described.

[0119] When the touch sensor 280K receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including information such as touch coordinates and the timestamp of the touch operation). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking the touch operation as a touch click operation and the control corresponding to the click operation being the control of the camera application icon as an example, the camera application calls the interface of the application framework layer to start the camera application, and then starts the camera driver by calling the kernel layer to capture a static image or video through the camera 293.

[0120] Taking the electronic device as a mobile phone as an example, the visual media search method provided by the embodiments of the present application will be introduced. The visual media search method provided by the embodiments of the present application can be applied to application programs such as the gallery application and the file management application.

[0121] Next, the interfaces and search logics involved in the visual media search method provided by the embodiments of the present application will be exemplarily described in combination with the accompanying drawings.

[0122] As Figure 3A shown in (a) of [], the interface 301 displays search history 303 and a "Clear" option 304. The search history 303 includes search statements that the user has entered, such as: "Watching the sunrise on the mountain top", "The sky taken in the photo". Other content displayed on the interface 301 can refer to the relevant display content of the above interface 105 and will not be elaborated here. Among them, the interface 301 can be called a search interface. The mobile phone can respond to the user'sFigure 1A Performing a click operation on the search box 104 in the interface 103 shown in (b) of displays the interface 301.

[0123] As Figure 3A shown in (b) of , in the interface 305 (i.e., the search result interface), the user enters the search statement "warming the tea around the stove" in the search box 306 of the interface 305, and the mobile phone searches for 239 pictures. The mobile phone displays some search results of the search statement "warming the tea around the stove" (e.g., thumbnails of 8 pictures) and the "More" option 307 corresponding to the search results of the search statement "warming the tea around the stove" on the interface 305. In response to the user's operation on the "More" option 307, the mobile phone displays the interface 308 shown in (c) of . The interface 308 can display the pictures in the search results of the search statement "warming the tea around the stove" in descending order according to the matching degree (or similarity) between the pictures and the search statement "warming the tea around the stove". The user can also perform an upward sliding operation on the interface 308 to view the pictures that have not been displayed. Figure 3A

[0124] Figure 3B In addition, the mobile phone can also be provided with a negative first screen, a pull-down search interface, etc. It can be understood that the negative first screen can be the leftmost split screen of the electronic device, which is used to provide functions such as search and quick services for the user. Among them, the negative first screen can also be used to display notification messages to be pushed to the user, such as application messages subscribed by the user, real-time hot search messages, segment selection, itinerary information, etc. The pull-down search interface is an interface displayed in response to the user's pull-down operation on the main interface. This interface is used to provide functions such as search and application suggestions for the user. This interface and the interface 315 in (c) of can be the same interface.

[0125]

[0126] Next, the negative first screen will be taken as an example for introduction. When the user needs to view the negative first screen of the mobile phone, the user can slide the screen of the mobile phone so that the electronic device displays the negative first screen. Figure 3B Figure 3B Exemplarily, as shown in (a) of , the mobile phone can receive the operation 1 performed by the user on the interface 309 (which can be called the desktop) of the mobile phone. Exemplarily, this operation 1 can be a rightward sliding operation as shown in (a) of . In response to this operation 1, the mobile phone can display the negative first screen 310 shown in (b) of . Among them, the negative first screen 310 can include: a search box 311, quick services 312, default cards 313, recommended cards 314, etc. The quick services 312 can be quick access to a certain page or function of an application program, such as: scan code, payment code, bus code, etc.; the default cards can be: gallery cards, remaining battery cards, etc.; the recommended cards can be recommended application cards. Figure 3B

[0127] ​​​​​The mobile phone receives a click operation by the user on the search box 311 on the minus one screen 310. In response to this click operation, the mobile phone can display the interface 315 shown in (c) of Figure 3B . The interface 315 may include: a search box 316 and application suggestions. The application suggestions include: icons of applications recommended for use. The interface 315 may also include: a search history 317 and its corresponding "clear" option 318. In response to a triggering operation by the user on the "clear" option 318, the search history 317 and the "clear" option 318 are no longer displayed on the interface 315. In addition, a hot news title, such as "Marathon race", may be displayed in the search box 316.

[0128] As shown in (d) of Figure 3B , in the interface 319, the search box of the interface 319 displays a search statement "Peaks photographed on weekends" input by the user. The interface 319 also displays a preview area 322 of the search results corresponding to the search statement "Peaks photographed on weekends" by the gallery application and a "Search in app" option 323 corresponding to the gallery application. In response to a triggering operation by the user on the preview area 322, the mobile phone can enter a photo details interface provided by the gallery application for the user to flip through the search results corresponding to the search statement "Peaks photographed on weekends", visual media matching the search statement, such as photos or videos. In response to a triggering operation by the user on the "Search in app" option 323, the mobile phone displays the interface 324 provided by the gallery application shown in (e) of Figure 3B . The interface 324 displays some search results of the search statement "Peaks photographed on weekends" and a "More" option corresponding to the search results of the search statement "Peaks photographed on weekends". In response to a triggering operation by the user on the "More" option, the mobile phone can display a search result details interface, and the search result details interface shows the pictures in the search results of the search statement "Peaks photographed on weekends". In addition, an online search option 321 may also be displayed in the interface 319. In response to a triggering operation by the user on the online search option 321, the mobile phone displays a search web page and shows online search results in the search web page.

[0129] The interfaces and search logics involved in the visual media search method are introduced above. Next, the specific implementation process of the visual media search method will be continued in combination with the software structure shown above in Figure 2B . As shown in Figure 4 , the implementation process may include S401 - S423. Among them, S401 - S407 may belong to the index construction stage, and S408 - S423 may belong to the search stage.

[0130] S401. The gallery service module adds and / or modifies visual media and their attributes.

[0131] The attributes of visual media may include one or more of the collection location, collection time, visual media name, face ID of the person in the visual media, person name, and the relationship between the person and the mobile phone user. Taking a captured video or picture as an example, the collection location refers to the shooting location, and the collection time refers to the shooting time; taking a screenshot as an example, the collection location refers to the screenshot location, and the collection time refers to the screenshot time; taking a downloaded video or picture as an example, the collection location refers to the download location, and the collection time refers to the download time.

[0132] The user can add visual media by means of shooting, downloading, screenshotting, etc. In addition, the user can also modify the existing visual media to obtain new visual media. The modifications include but are not limited to operations such as beautification, custom naming, and adding watermarks.

[0133] S402. The gallery service module stores the visual media and its attributes.

[0134] The gallery service module can, in response to the addition or modification operation of the visual media input by the user, store the visual media and its attributes locally on the mobile phone. In practical applications, with the authorization of the user, the mobile phone can store the locally stored visual media and its attributes in the cloud to relieve the storage pressure on the local mobile phone.

[0135] S403. The gallery service module sends Request 1 to the multimodal understanding module. Among them, Request 1 is used to trigger the visual semantic understanding of the visual media.

[0136] S404. The multimodal understanding module returns the image feature vector of the visual media to the gallery service module.

[0137] In the embodiment of the present application, the multimodal understanding module, in response to the above Request 1, determines the image feature vector of the visual media stored in the mobile phone. Since visual semantic understanding requires a large amount of computing resources, in order not to affect the user's use, when the mobile phone is in the charging and screen-off state, the gallery service module can request the multimodal understanding module to perform visual semantic understanding on the visual media stored locally on the mobile phone, such as performing visual semantic understanding on the newly added or modified visual media to obtain the visual semantic vector (or called the image feature vector) of the visual media, so as to realize the offline processing of the visual media and reduce the impact of the visual semantic understanding of the visual media on other services running on the mobile phone.

[0138] Among them, the multimodal understanding module can perform visual semantic understanding on visual media based on a multimodal model to obtain the image feature vector of the visual media. In addition, the multimodal model can be used not only for: performing visual semantic understanding on visual media to obtain the visual semantic vector of the visual media; but also for: performing semantic understanding on the search statement to obtain the text feature vector (or called sentence semantic vector) of the search statement. Among them, the process of the multimodal model performing semantic understanding on the search statement can refer to the relevant descriptions below and will not be introduced in detail here.

[0139] In some embodiments, the multimodal model can map visual media and text into vectors of the same dimension, that is to say, the dimension of the visual semantic vector of visual media is the same as the dimension of the semantic vector of text (for example: the sentence semantic vector of the search statement). The multimodal model can specifically be a contrastive language-image pre-training (CLIP) model. The mobile phone can map visual media and text into a unified vector space through the CLIP model to understand the relationship between different modal resources visually and textually, and then use it for image retrieval.

[0140] Among them, the CLIP model is a standard CLIP model, that is, an existing CLIP model, or the CLIP model is a custom CLIP model.

[0141] Exemplarily, the above custom CLIP model can include a base model and a fine-tuning model. Among them, the fine-tuning model is obtained by continuing to train the base model with a training set of a specific scenario on the basis of the base model. Therefore, the fine-tuning model has a support scope and can support the processing of images and their texts in a specific scenario. The generalization ability of the fine-tuning model is less than that of the base model, while the retrieval ability of the base model in a specific scenario is less than that of the fine-tuning model.

[0142] In some embodiments, the above-mentioned base model and fine-tuning model reuse part of the network to reduce the resource occupancy of the mobile phone installed with the custom CLIP model. The base model and the fine-tuning model need to process not only visual media but also text. The module for processing visual media in the base model can be called the visual encoding module, and the module for processing visual media in the fine-tuning model can be called the visual fine-tuning module. The above-mentioned image feature vectors can include the image feature vectors corresponding to the base model and the image feature vectors corresponding to the fine-tuning model. The visual encoding module of the base model and the visual fine-tuning module of the fine-tuning model can reuse the image encoder, so that the visual encoding module can use the feature vectors of the visual media output by the image encoder to continue to determine the first L2 norm of the visual media, and realize the determination of the image feature vectors corresponding to the base model. The visual fine-tuning module can use the feature vectors of the visual media output by the image encoder to continue to determine the second L2 norm of the visual media, and realize the determination of the image feature vectors corresponding to the fine-tuning model.

[0143] Exemplarily, after receiving Request 1 sent by the gallery service module, in response to this Request 1, as Figure 5A shown, the multimodal understanding module can input k visual media in the mobile phone into the visual encoding module. The image encoder in the visual encoding module encodes each visual media among the k visual media to obtain the image feature vector 1 of each visual media. The image feature vector of this visual media can be 768-dimensional, that is, X = {x1, x2,..., x768}, and X represents this image feature vector 1. After that, on the one hand, the visual encoding module can output this 768-dimensional image feature vector 1 so that this image feature vector 1 continues to be used as the input parameter of the visual fine-tuning module. On the other hand, the visual encoding module continues to use the mapping matrix 1 and the image feature vector 1 to calculate and output the first L2 norm α1 of the image feature vectors of each visual media among the k visual media. Among them, the dimension of the mapping matrix 1 is 768*512-dimensional to map the image feature vector 1 from 768 dimensions to 512 dimensions. Here, the image feature vector 1 of the visual media and the first L2 norm of the visual media can be regarded as the image feature vectors of the visual media corresponding to the base model.

[0144] Specifically, for each of the k visual media, the visual encoding module can use to calculate the first L2 norm α1 of the image feature vector of the visual media. Among them, X represents the 768-dimensional image feature vector 1, M 1 represents the mapping matrix 1, V 1 represents the image feature vector 2, and the dimension of V 1 is 512-dimensional, represents the i-th element in V 1 .

[0145] As Figure 5BAs shown, after the visual fine-tuning module receives the image feature vector 1 of each of the k visual media, for each visual media, it continues to use the mapping matrix 2 and the image feature vector 1 of this visual media to calculate and output the second L2 norm α2 of the image feature vector of this visual media. Among them, the dimension of the mapping matrix 2 is 768*512, so as to map the image feature vector 1 from 768 dimensions to 512 dimensions. Here, the image feature vector 1 of the visual media and the second L2 norm of the visual media can be regarded as the image feature vector of the visual media corresponding to the fine-tuning model.

[0146] Specifically, for each of the k visual media, the visual encoding module can use to calculate the second L2 norm α2 of the image feature vector of the visual media. Among them, X represents the 768-dimensional image feature vector 1, M 2 represents the mapping matrix 2, V 2 represents the image feature vector 3, and the dimension of V 2 is 512 dimensions, represents the i-th element in V 2 .

[0147] In some embodiments, the above fine-tuning model (such as the mapping matrix 2 in the above visual fine-tuning module) is obtained by continuing to train using a training set of a specific scenario. When the device trains the fine-tuning model, it can freeze the image encoder and only train the last fully connected layer M 2 of the fine-tuning model. The training set of this specific scenario can be an image training set including visual content corresponding to a specific whitelist. This specific whitelist represents the support range of the fine-tuning model. Among them, the device can be the above mobile phone or not. The present application does not limit the device for training the fine-tuning model.

[0148] After obtaining the image feature vector of the visual media (such as the image feature vector of the visual media corresponding to the above base model, the image feature vector of the visual media corresponding to the fine-tuning model), in order to facilitate retrieving the visual media required by the user using the image feature vector of the visual media, the multimedia understanding module can save the image feature vector of the visual media, that is, the above image feature vector 1, the first L2 norm α1 and the second L2 norm α2, so that when saving the image feature vector of a visual media, only a 770-dimensional visual vector needs to be saved, instead of saving 512*2 dimensions, that is, 1024-dimensional visual features, reducing the resources required to store the image feature vector. Among them, the 512-dimensional vector is the image feature vector obtained by processing the 768-dimensional vector using the mapping matrix 1 or the mapping matrix 2 (that is, the above V 1 and V 2 ).

[0149] It should be noted that the above image encoder located in the visual encoding module is only an example. The image encoder can also be located in the visual fine-tuning module, and the present application does not limit it. Additionally, the image feature vectors of the visual media mentioned above, which include the image feature vectors of the visual media corresponding to the base model and the image feature vectors of the visual media corresponding to the fine-tuning model, are only an example. The image feature vectors of the visual media can also only include the image feature vectors determined by one model, and this model can be a clip model or not a clip model.

[0150] As introduced above, the determination of the image feature vectors of the visual media in the mobile phone can be done offline, while the text feature vectors of the search query can be determined online after the mobile phone receives the search query input by the user, enabling the mobile phone to utilize the text feature vectors of the search query and the image feature vectors of the visual media in the mobile phone to determine the visual media that matches the search query. Among them, the process of determining the text feature vectors of the search query and using the text feature vectors of the search query and the image feature vectors of the visual media to determine the visual media that matches the search query can be referred to below, and will not be introduced in detail here. Next, the process of establishing the index of the visual media will be introduced first.

[0151] S405. The gallery service module stores the image feature vectors of the visual media.

[0152] S406. The gallery service module sends the attribute information and visual semantic vectors of the visual media to the search module.

[0153] In the embodiments of the present application, the gallery service module can store the image feature vectors of each of the k visual media locally in the mobile phone. Moreover, the gallery service module can send the attribute information and image feature vectors of the visual media to the search module in batches for the search module to construct the index of the visual media.

[0154] Among them, optionally, the multi-gallery service module may not upload the image feature vectors of the visual media to the cloud. Of course, it can also upload the image feature vectors of the visual media to the cloud under user authorization, and the present application does not limit it.

[0155] S407. The search module constructs the index of the visual media.

[0156] The index of the visual media constructed by the search module may include: the attributes of the visual media and / or the visual semantic vectors of the visual media.

[0157] S408. The gallery service module receives the search query input by the user.

[0158] The user can use the search interface provided by the gallery service module. For example, the user can Figure 3AEnter "warming the tea around the stove" in the search box 306 on the interface 305 shown in (b) of []. This "warming the tea around the stove" is the search statement.

[0159] S409. The gallery service module sends the search statement to the search module.

[0160] S410. The search module determines whether the search statement includes visual content.

[0161] In the embodiments of the present application, the search module can determine whether the search statement includes visual content, that is, determine whether it is necessary to search for visual media using image feature vectors and text feature vectors. In other words, the search module determines whether it is necessary to use a multimodal model to determine visual media.

[0162] In the case where the search statement does not include visual content, it indicates that visual media can be searched using the attributes of visual media, and there is no need to search for visual media using image feature vectors and text feature vectors. That is, it indicates that there is no need to use a multimodal model to determine visual media, and the search module can execute S411.

[0163] In the case where the search statement includes visual content, it indicates that it is necessary to search for visual media using image feature vectors and text feature vectors. That is, it indicates that it is necessary to use a multimodal model to determine visual media, and the search module can execute S412.

[0164] In some embodiments, the search module may determine whether a search statement includes visual content by determining whether the search statement includes a visual semantic subject. The search module may send Request 2 to the natural language understanding module. The natural language understanding module may perform semantic subject recognition on the search statement to obtain the semantic subjects included in the search statement. Among them, the natural language understanding module performs semantic subject recognition based on a natural language understanding model. Specifically, the natural language understanding module may use named entity recognition (NER) technology to perform semantic subject recognition on the search statement to obtain the semantic subjects included in the search statement. In the embodiments of the present application, the semantic subject may also be referred to as an entity, and the semantic subject may include one or more of: a semantic subject related to time, a semantic subject related to location, a semantic subject related to a person's name, a semantic subject related to a person's relationship, and a semantic subject related to visual content (or referred to as a visual semantic subject). Among them, optionally, the semantic subject related to visual content is determined from M (M≥1) preset semantic subjects related to visual content through named entity recognition technology. These M semantic subjects can be configured by developers of the gallery service module according to actual situations. Generally, these M semantic subjects are all nouns. Exemplarily, for the search statement "The sky photographed in City 1 on September 1st", semantic subject recognition can obtain three semantic subjects: "September 1st", "City 1", and "the sky". Among them, "the sky" is related to visual content and can be a visual semantic subject.

[0165] After that, the natural language understanding module may return the recognized semantic subjects to the search module. After that, the search module may determine whether the semantic subjects include a visual semantic subject. In the case where the semantic subjects do not include a visual semantic subject, the search module may determine that the search statement does not include visual content. In the case where the semantic subjects include a visual semantic subject, the search module may determine that the search statement includes visual content.

[0166] S411. The search module performs recall based on the index of the visual media and the search statement to obtain search results.

[0167] Exemplarily, when the search statement does not include a visual semantic entity, it indicates that the search statement does not include content related to visual semantics. The search module can directly query the visual media corresponding to the semantic entity based on the constructed index, that is, based on the attributes of each visual media, and use it as the search result. Here, the semantic entity represents the attribute of the visual media. For example, if the search statement is "September 1st" without visual content, the search module can, based on the index, find the visual media with the acquisition time of September 1st to obtain the search result. Another example is that if the search statement includes "a photo of Zhang San" without visual content, the search module can match the names of each visual media with "Zhang San" to determine the visual media with the name attribute of "Zhang San" and use it as the search result.

[0168] S412. The search module filters out the non-visual semantic entities in the above search statement to obtain a filtered search statement.

[0169] Exemplarily, when the above semantic entity includes a visual semantic entity, it indicates that the search statement includes content related to visual semantics and may include content unnecessary for visual semantics, that is, non-visual semantic entities. Since non-visual semantic entities are irrelevant to the search for visual content, the search module can first filter out the non-visual semantic entities in the search statement to obtain a filtered search statement for searching visual media that matches the filtered search statement.

[0170] Among them, optionally, the non-visual semantic entity refers to: semantic entities related to time, semantic entities related to location, semantic entities related to personal names, semantic entities related to personal relationships, etc., which are semantic entities related to the attributes of visual media. It should be understood that since personal relationships have been regarded as attributes of visual media before, personal relationships can be regarded as non-semantic entities here.

[0171] For example, for the search statement "the sky photographed this year", "this year" is a non-visual semantic entity. Correspondingly, the filtered search statement is "the sky photographed".

[0172] In practical applications, after filtering out semantic entities irrelevant to visual content, there may be some redundant stop words. For example, for the search statement "the sky photographed in City 1 this year", after deleting "this year" and "City 1", the stop word "in" becomes a redundant word, so the search module can also filter it out. Specifically, the search module can filter out the non-visual semantic entities and their related stop words in the search statement to obtain a filtered search statement. For example, the filtered search statement corresponding to the search statement "the sky photographed in City 1 this year" is "the sky photographed".

[0173] S413. The search module sends the filtered search statement to the multimodal understanding module.

[0174] S414. The multi-modal understanding module performs semantic understanding on the filtered search statement to obtain the text feature vector of the filtered search statement.

[0175] Exemplarily, the multi-modal understanding module may use a multi-modal model to determine the text feature vector of the filtered search statement. Optionally, the multi-modal model may be a CLIP model. The CLIP model may be a standard CLIP model, or the CLIP module is a custom CLIP model.

[0176] In some embodiments, from the above, it can be seen that the above image feature vector may include the image feature vector corresponding to the base model and the image feature vector corresponding to the fine-tuning model. The visual encoding module of the base model and the visual fine-tuning module of the fine-tuning model may reuse the image encoder. Correspondingly, in order to keep the similarity of the text-image pair consistent, the present application introduces a text encoding module, so that the text encoding module can reuse the text encoder to respectively output the text feature vector of the filtered search statement corresponding to the base model and the text feature vector of the filtered search statement corresponding to the fine-tuning model, for calculating the similarity between the search statement and the visual media by using the text feature vector of the filtered search statement corresponding to the base model and the image feature vector of the visual media corresponding to the base model, and calculating the similarity between the search statement and the visual media by using the text feature vector of the filtered search statement corresponding to the fine-tuning model and the image feature vector of the visual media corresponding to the fine-tuning model, to realize the calculation of the similarity of the text-image pair.

[0177] Exemplarily, as Figure 5C shown, the multi-modal understanding model inputs the filtered search statement into the CLIP model. The text encoder in the text encoding module of the CLIP model encodes the filtered search statement to obtain and output the feature vector of the filtered search statement, and the dimension of the feature vector is 768. Then, the text encoding module may input the feature vector of the filtered search statement into the base model branch and the fine-tuning model branch respectively. Among them, the structures of the base model branch and the fine-tuning model branch are similar.

[0178] For the base model branch: The text encoding module uses the mapping matrix 3 and the feature vector of the filtered search statement to obtain the intermediate variable 1 with a dimension of 512. Exemplarily, the text encoding module calculates the intermediate variable 1 according to T 1 =YN 1 , where T 1 represents the intermediate variable 1, Y represents the feature vector of the filtered search statement, and N 1 represents the mapping matrix 3. The text encoding module calculates the first L2 norm of the filtered search statement by using the intermediate variable 1. Exemplarily, the text encoding module may calculate the first L2 norm of the filtered search statement according to . Among them, β 1 represents the first L2 norm of the filtered search statement. Represents the i-th element in T1.

[0179] Since the dimension of the image feature vector of the visual media is 768-dimensional, and in order to perform calculations between the image feature vector and the text feature vector, while the intermediate variable 1 of the filtered search statement is 512-dimensional, therefore, the text encoding module needs to map the intermediate variable 1 of the filtered search statement to 768 dimensions. Then, the text encoding module can use the above mapping matrix 1 and the first L2 norm of the filtered search statement to map the intermediate variable 1 into a 768-dimensional text feature vector, so as to obtain the text feature vector of the filtered search statement corresponding to the base model. Specifically, the text encoding module can use To calculate the text feature vector of the filtered search statement corresponding to the base model. Among them, T' 1 Represents the text feature vector of the filtered search statement corresponding to the base model, Is the transpose matrix of M 1 And M 1 Represents the above mapping matrix 1.

[0180] For the fine-tuning model branch: The text encoding module uses the mapping matrix 4 and the feature vector of the filtered search statement to obtain the intermediate variable 2 of 512 dimensions. Exemplarily, the text encoding module calculates the intermediate variable 2 according to T 2 =YN 2 , where T 2 Represents the intermediate variable 2, N 2 Represents the mapping matrix 4, and the text encoding module calculates the second L2 norm of the filtered search statement using the intermediate variable 2. Exemplarily, the text encoding module can calculate according to To calculate the second L2 norm of the filtered search statement. Among them, β 2 Represents the second L2 norm of the filtered search statement, Represents the i-th element in T 2 In.

[0181] Since the dimension of the image feature vector of the visual media is 768-dimensional, and in order to perform calculations between the image feature vector and the text feature vector, while the intermediate variable 2 of the filtered search statement is 512-dimensional, therefore, the text encoding module needs to map the intermediate variable 2 of the filtered search statement to 768 dimensions. Then, the text encoding module can use the above mapping matrix 2 and the first L2 norm of the filtered search statement to map the intermediate variable 2 into a 768-dimensional text feature vector, so as to obtain the text feature vector of the filtered search statement corresponding to the fine-tuning model. Specifically, the text encoding module can use To calculate the text feature vector of the filtered search statement corresponding to the base model. Among them, T' 2 Represents the text feature vector of the filtered search statement corresponding to the fine-tuning model, Is M 2The transposed matrix of, M 1 represents the above mapping matrix 2.

[0182] It should be noted that the text feature vector of the filtered search statement corresponding to the fine-tuning model and the text feature vector of the filtered search statement corresponding to the base model can be output by the text encoding module simultaneously.

[0183] In the embodiments of the present application, the base model and the fine-tuning model in the custom CLIP model reuse the text encoder, so that the custom CLIP model only needs to calculate the feature vector of the filtered search statement once, and then can use the feature vector of the filtered search statement to respectively determine the text feature vector of the filtered search statement corresponding to the base model and the text feature vector of the filtered search statement corresponding to the fine-tuning model, without the base model and the fine-tuning model calculating the feature vector of the filtered search statement separately, thereby improving the calculation efficiency of the text feature vector of the filtered search statement.

[0184] In the embodiments of the present application, in order to ensure the generalization ability of the multi-modal model and the ability to recognize specific images and texts, the present application sets the multi-modal model to include a base model and a fine-tuning model. Since the base model and the fine-tuning model have the same model weights, that is, part of the network is the same, directly setting two models will cause waste of mobile phone memory. Therefore, the present application reuses the image encoder and the text encoder for the base model and the fine-tuning model, so that the base model and the fine-tuning model can respectively use the feature vectors output by the image encoder and the text encoder to determine their respective corresponding text feature vectors and image feature vectors. Generally speaking, the determination time of the image feature vector and the text feature vector is reduced by nearly half.

[0185] S415. The multi-modal understanding module performs a recall based on the text feature vector and the image feature vector of the visual media to obtain candidate visual media.

[0186] In some embodiments, the image feature vector of the above visual media may be sent by the search module to the multi-modal understanding module, or may be obtained by the multi-modal understanding module from the local of the mobile phone.

[0187] In the embodiments of the present application, for each visual media on the mobile phone (such as each visual media among the above k visual media), the multi-modal understanding module can calculate the vector similarity (or simply referred to as similarity) between the image feature vector of the visual media and the text feature vector of the filtered search statement, that is, calculate the similarity between the visual media and the filtered search statement based on the image feature vector of the visual media and the text feature vector of the filtered search statement. Then, the multi-modal understanding module determines the visual media that matches the filtered search statement from the k visual media according to the vector similarity, and uses the determined visual media as the candidate visual media.

[0188] Exemplarily, the multimodal understanding module may determine visual media with a vector similarity greater than or equal to threshold 1 as visual media matching the filtered search statement. Optionally, in a case where the number of visual media with a vector similarity greater than or equal to threshold 1 is greater than quantity 1, the multimodal understanding module may sort the visual media with a vector similarity greater than or equal to threshold 1 in descending order of vector similarity. Subsequently, the multimodal understanding module may determine the top n visual media as visual media matching the filtered search statement. Here, n is a positive integer.

[0189] In some embodiments, the image feature vectors of the above visual media may include the image feature vectors of the visual media corresponding to the base model and the image feature vectors of the visual media corresponding to the fine-tuning model. The text feature vectors of the above filtered search statement include the text feature vectors of the filtered search statement corresponding to the base model and the text feature vectors of the filtered search statement corresponding to the fine-tuning model. In one case, as Figure 5D shown, the multimodal understanding module may calculate the vector similarity 1 between the text feature vector of the filtered search statement corresponding to the base model and the image feature vector of the visual media corresponding to the base model (i.e., calculate the similarity between the visual media corresponding to the base model and the filtered search statement), and calculate the vector similarity 2 between the text feature vector of the filtered search statement corresponding to the fine-tuning model and the image feature vector of the visual media corresponding to the fine-tuning model (i.e., calculate the similarity between the visual media corresponding to the fine-tuning model and the filtered search statement). Subsequently, the multimodal understanding module may determine visual media with vector similarity 1 or vector similarity 2 greater than or equal to threshold 1 as visual media matching the filtered search statement.

[0190] In another case, the multimodal understanding module may use whether the filtered search statement hits whitelist 1 to determine whether the filtered search statement is within the support range of the fine-tuning model, that is, determine whether to use the fine-tuning model branch to determine candidate visual media. The process of the multimodal understanding module using whitelist 1 to determine candidate visual media will be described below in conjunction with Figure 6 , to introduce the process of the multimodal understanding module using whitelist 1 to determine candidate visual media.

[0191] S501. The multimodal understanding module determines whether the filtered search statement belongs to whitelist 1.

[0192] In the embodiments of the present application, when the multimodal understanding module determines that the filtered search statement does not belong to the whitelist, it indicates that the filtered search statement is within the support range of the base model. The multimodal understanding module may use the base model branch to determine candidate visual media, and the multimodal understanding module may execute S502.

[0193] When the filtered search statement belongs to the whitelist, it indicates that the search statement is within the support scope of the fine-tuning model. The multimodal understanding module can enable the fine-tuning model branch to determine candidate visual media, and the multimodal understanding module can execute S504.

[0194] In some embodiments, the multimodal understanding module determines whether each semantic entity (i.e., visual semantic entity) in the filtered search statement belongs to Whitelist 1. Considering that the search statement input by the user is generally a phrase, the multimodal understanding module can determine that the filtered search statement belongs to the whitelist when all semantic entities in the filtered search statement belong to Whitelist 1. When there is a semantic entity that does not belong to Whitelist 1, it is determined that the filtered search statement does not belong to the whitelist. For example, the filtered search statement is "a boy holding a basket", and the semantic entities include "basket" and "boy". The multimodal understanding module can respectively determine whether the basket and the boy belong to Whitelist 1. When both the basket and the boy belong to Whitelist 1, the multimodal understanding module can determine that the filtered search statement belongs to the whitelist. When either the basket or the boy does not belong to Whitelist 1, the multimodal understanding module can determine that the filtered search statement does not belong to the whitelist.

[0195] S502. For each visual media, the multimodal understanding module calculates the vector similarity 1 between the image feature vector of the visual media corresponding to the base model and the text feature vector of the filtered search statement corresponding to the base model.

[0196] In the embodiments of the present application, as Figure 7 shown, when the filtered search statement does not belong to Whitelist 1, the multimodal understanding module can call the text feature vector of the filtered search statement output by the base model branch of the text encoding module and the image feature vector of the visual media corresponding to the base model, and calculate the score corresponding to the visual media corresponding to the base model, that is, calculate the vector similarity 1 between the visual media corresponding to the base model and the filtered search statement.

[0197] Specifically, the multimodal understanding module can pass to calculate the score corresponding to the visual media corresponding to the base model. Among them, S 1 represents the score corresponding to the visual media corresponding to the base model, α1 represents the first L2 norm of the above-mentioned image feature vector of the visual media. X represents the above-mentioned image feature vector 1, and T' 1 represents the text feature vector of the filtered search statement corresponding to the above-mentioned base model.

[0198] S503. The multimodal understanding module uses the visual media with a vector similarity 1 greater than the threshold 1 as the candidate visual media.

[0199] S504. For each visual medium, the multimodal understanding module calculates the vector similarity 2 between the image feature vector of the visual medium corresponding to the fine-tuning model and the text feature vector of the filtered search statement corresponding to the fine-tuning model.

[0200] In the embodiments of the present application, as described above Figure 7 As shown, when the filtered search statement belongs to the whitelist 1, the multimodal understanding module can call the text feature vector of the filtered search statement output by the fine-tuning model branch of the text encoding module and the image feature vector of the visual medium corresponding to the fine-tuning model, and calculate the score corresponding to the visual medium corresponding to the fine-tuning model, that is, calculate the vector similarity 2 between the visual medium corresponding to the fine-tuning model and the filtered search statement.

[0201] Specifically, the multimodal understanding module can pass through Calculate the score corresponding to the visual medium corresponding to the base model. Among them, S 2 Represents the score corresponding to the visual medium corresponding to the fine-tuning model, and α 2 Represents the second L2 norm of the image feature vector of the above visual medium. X represents the above image feature vector 1, and T' 2 Represents the text feature vector of the filtered search statement corresponding to the above fine-tuning model.

[0202] S505. The multimodal understanding module uses the visual media with a vector similarity 2 greater than the threshold 1 as candidate visual media.

[0203] In some embodiments, the operations performed by the above multimodal understanding module can be performed by the multimodal model. In addition, the above S501-S505 can be performed by the output module of the multimodal understanding module, that is, the model in the multimodal model.

[0204] In the embodiments of the present application, the multimodal understanding module can determine the image feature vector of the visual medium offline, and only needs to determine the text feature vector of the search statement online, which improves the calculation time of the vector similarity between the image feature vector and the text feature vector, thereby effectively reducing the user's retrieval time. In addition, the multimodal understanding module can determine whether to use the fine-tuning model to determine the candidate visual media according to whether the filtered search statement belongs to the whitelist 1, so as to improve the determination accuracy of the candidate visual media, thereby improving the accuracy of the search results.

[0205] Optionally, the above whitelist 1 can be extended, that is, the support range of the fine-tuning model can be extended. In other words, the scenarios supported by the fine-tuning model can be extended. It only needs to train the fine-tuning model with the training set corresponding to the extended scenario. It should be understood that training the fine-tuning model is actually training the fine-tuning model branch in the above fine-tuning module and text encoding module, without training the reused parts (such as the above image encoder, text encoder).

[0206] In some embodiments, after determining the candidate visual media, the multimodal understanding module can directly use the above candidate visual media as the search result. After that, the multimodal understanding module can return the search result to the search module. Then, the search module can send the search result to the gallery service module for the gallery service module to display the search result.

[0207] In addition, to improve the accuracy of the search result, after obtaining the above candidate visual media, the mobile phone can continue to perform a secondary confirmation on the candidate visual media to further screen visual media from the candidate visual media to obtain visual media that matches the filtered search statement, that is, the search statement. Optionally, the mobile phone can use both the visual media with vector similarity 1 greater than threshold 1 and the visual media with vector similarity 2 greater than threshold 1 as candidate visual media. In other words, the mobile phone can use the visual media determined by the base model and the fine-tuned model respectively as candidate visual media, that is, perform two searches. The process of secondary confirmation of the candidate visual media is introduced below.

[0208] S416. The multimodal understanding module sends the information of the filtered search statement and the image information of the candidate visual media to the secondary confirmation module.

[0209] The information of the filtered search statement includes one or more of the filtered search statement, the word segmentation result of the filtered search statement, label 1 included in the filtered search statement, and the text feature vector of the filtered search statement.

[0210] The image information of the candidate visual media includes one or more of the similarity between the candidate visual media and the filtered search statement, label 2 of the candidate visual media, and the image feature vector of the candidate visual media.

[0211] In some embodiments, the above image feature vector and text feature vector are determined based on a multimodal model including a base model and a fine-tuned model. Correspondingly, the text feature vector of the filtered search statement can include the text feature vector of the filtered search statement corresponding to the base model and the text feature vector of the filtered search statement corresponding to the fine-tuned model. The image feature vector of the candidate visual media can include the image feature vector of the candidate visual media corresponding to the base model and the image feature vector of the candidate visual media corresponding to the fine-tuned model.

[0212] The similarity between the above candidate visual media and the filtered search statement can include the similarity between the candidate visual media corresponding to the base model and the filtered search statement (or referred to as similarity 1), and the similarity between the candidate visual media corresponding to the fine-tuned model and the filtered search statement (or referred to as similarity 2).

[0213] Of course, the above image feature vectors, text feature vectors, and similarity can also be determined by only one model (such as the above base model or fine-tuned model), and the present application does not limit it.

[0214] In some embodiments, the multimodal understanding module can perform word segmentation extraction on the filtered search statement to obtain the word segmentation result of the filtered search statement. For example, if the filtered search statement is "The child is playing by the sea", the word segmentation result is: child, is, by the sea, and playing. Exemplarily, the multimodal understanding module can perform word segmentation extraction on the filtered search statement through the natural language understanding engine service (the natural language understanding, NLU). Optionally, the multimodal understanding module can perform word segmentation extraction on the filtered search statement through the natural language understanding module.

[0215] After obtaining the word segmentation result of the filtered search statement, the multimodal understanding module can perform label mapping on the word segmentation result to obtain the label included in the filtered search statement, that is, label 1. Specifically, for each word in the word segmentation result, the filtering search module determines whether the word exists in the preset label table. In the case where the word does not exist in the preset label table, the multimodal understanding module can determine that the word does not have a corresponding label, that is, the filtered search statement does not include the label corresponding to the word.

[0216] In the case where the word exists in the preset label table, the multimodal understanding module can use the label corresponding to the word in the preset label table as the label 1 included in the filtered search statement. For example, the word segmentation includes "cat", and the label corresponding to "cat" in the preset label table is "cat", so the label included in the filtered search statement includes "cat". It should be noted that the words in the word segmentation result of the filtered search statement and the labels corresponding to the words may be the same or different.

[0217] In some embodiments, the above label 2 of the candidate visual media is obtained from the label library. The label 2 of the visual media in the label library represents the classification label of the visual media, which can be determined by the mobile phone (such as the multimodal understanding module in the mobile phone) using the image classification model to identify the visual media. Exemplarily, the classification label of the visual media can also be determined offline by the mobile phone.

[0218] S417. The secondary confirmation module determines the Boolean value corresponding to each candidate visual media based on the information of the filtered search statement and the image information of the candidate visual media.

[0219] S418. For each candidate visual media, when the Boolean value corresponding to the candidate visual media is true, the secondary confirmation module determines that the candidate visual media is visual media 1.

[0220] In the embodiments of the present application, the multimodal understanding module inputs the information of the filtered search statement and the image information of the candidate visual media into the secondary confirmation module, so that the secondary confirmation module determines the Boolean value corresponding to each candidate visual media, and the Boolean value corresponding to the candidate visual media indicates whether the candidate visual media is a visual media matching the filtered search statement, realizing the secondary confirmation of the candidate visual media.

[0221] In some embodiments, the above-mentioned secondary confirmation module may include a binary classification module, a dynamic threshold module, a label confirmation module, and a whitelist threshold module. The binary classification module may adopt a binary classification model to determine whether there is a correlation between the filtered search statement and the candidate visual media, so as to obtain the Boolean value of the candidate visual media. The dynamic threshold module may adopt a dynamic threshold model to determine the threshold 2 matching the length of the filtered search statement, so that the similarity between the candidate visual media and the filtered search statement can be compared with the threshold 2 to obtain the Boolean value of the candidate visual media. The label confirmation module may compare the label 2 of the candidate visual media with the label 1 included in the filtered search statement to obtain the Boolean value of the candidate visual media. The whitelist threshold module may determine the threshold 3 that matches the filtered search statement as a whole according to whether the filtered search statement hits the preset dictionary, so that the similarity between the candidate visual media and the filtered search statement can be compared with the threshold 3 to obtain the Boolean value of the candidate visual media. Among them, the detailed process of the secondary determination module determining the candidate visual media through the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module, that is, the implementation process of S417 above, can refer to the relevant descriptions below and will not be introduced here first.

[0222] S419. The secondary confirmation module returns the visual media 1 to the search module.

[0223] S420. The search module obtains the visual media 2 based on the non-visual semantic entity in the above search statement and the index of the visual media.

[0224] S421. The search module takes the intersection of the visual media 1 and the visual media 2 as the search result.

[0225] In the embodiments of the present application, the search module searches for the visual media (or referred to as visual media 2) corresponding to the filtered non-visual semantic entity in the search statement input by the user from the constructed index, that is, searches for the visual media whose attributes match the non-visual semantic entity, and takes it as the visual media 2.

[0226] After that, the search module can determine the intersection of Visual Media 2 and Visual Media 1 to obtain the target visual media. The attributes of the target visual media match the non-visual semantic subject of the search statement, and the visual content of the target search media corresponds to the visual semantic subject of the search statement. For example, if the search statement is "the sky photographed this year", then the above Visual Media 1 includes sky content, and the acquisition time of Visual Media 2 is this year. In this case, the target visual media includes sky content, and the acquisition time of the target visual media is this year.

[0227] S422. The search module sends the search results to the gallery service module.

[0228] S423. The gallery service module displays the search results.

[0229] In the embodiments of the present application, the gallery service module can display the search results to the user. For example, the search results can be displayed through Interface 305 above Figure 3A and Interface 308 above Figure 3A . As shown in Interface 305 above Figure 3A , the search results include 239 pictures, and Interface 305 only displays the thumbnails of 8 pictures in the search results. If the user wants to view these 239 pictures, the user can click the "More" option in Interface 305. In response to this operation, the mobile phone displays Interface 308 as shown in Figure 3A .

[0230] In some embodiments, the gallery service module can sort the target visual media in the search results according to the acquisition time and display the target visual media. For example, the target visual media are sorted in ascending order of acquisition time, so that the earlier the acquisition time, the higher the display order.

[0231] In other embodiments, the gallery service module can sort the target visual media in the search results according to the similarity between the target visual media and the filtered search statement. For example, the target visual media are sorted in descending order of similarity, so that the higher the similarity of the target visual media, the higher the display order. That is to say, the similarity between the target visual media displayed earlier and the filtered search statement is greater than or equal to the similarity between the target visual media displayed later and the filtered search statement. Here, the similarity between the target visual media and the filtered search statement can be Similarity 1 or Similarity 2, or a similarity calculated based on Similarity 1 and Similarity 2.

[0232] Optionally, when the filtered search statement belongs to the above-mentioned whitelist 1, the similarity between the target visual media and the filtered search statement can be similarity 2, that is, the similarity between the visual media corresponding to the fine-tuning model and the filtered search statement. When the filtered search statement does not belong to the above-mentioned whitelist 1, the similarity between the target visual media and the filtered search statement can be similarity 1, that is, the similarity between the visual media corresponding to the base model and the filtered search statement.

[0233] Optionally, the similarity calculated based on similarity 1 and similarity 2 can be the average of similarity 1 and similarity 2. Alternatively, calculate the weighted sum of similarity 1 and similarity 2, and this application does not limit it.

[0234] Next, a possible implementation process of the above S417 will be continued, as Figure 8 shown, this process may include S417a - S417g.

[0235] S417a. The secondary confirmation module determines whether the filtered search statement exists in the preset fine-tuning model vocabulary.

[0236] In the embodiments of the present application, the secondary confirmation module can determine whether the filtered search statement exists in the preset fine-tuning model vocabulary to determine whether the filtered search statement corresponds to the branch of the base model or the branch of the fine-tuning model, that is, to determine whether to perform secondary confirmation through the module corresponding to the base model or the module corresponding to the fine-tuning model.

[0237] When the filtered search statement does not exist in the above-mentioned preset fine-tuning vocabulary, it indicates that the filtered search statement corresponds to the branch of the base model, that is, it indicates that the candidate visual media is secondarily confirmed through the module corresponding to the base model to select a visual media that better matches the filtered search statement from the candidate visual media, and the secondary confirmation module can execute S417b.

[0238] When the filtered search statement exists in the above-mentioned preset fine-tuning vocabulary, it indicates that the filtered search statement corresponds to the branch of the fine-tuning model, that is, it indicates that the candidate visual media is secondarily confirmed through the module corresponding to the fine-tuning model to select a visual media that better matches the filtered search statement from the candidate visual media, and the secondary confirmation module can execute S417f.

[0239] S417b. The secondary confirmation module inputs the text feature vector of the filtered search statement and the image feature vectors of each candidate visual media into the binary classification module to obtain the Boolean value 1 corresponding to each candidate visual media output by the binary classification module.

[0240] In the embodiments of the present application, it is assumed that the number of candidate visual media is m. As Figure 9AAs shown in the figure, the secondary confirmation module inputs the image feature vectors of each candidate visual medium among the m candidate visual media and the text feature vector of the filtered search statement into the binary classification model in the binary classification module. For each candidate visual medium among the m candidate visual media, the binary classification model determines the matching degree between the candidate visual medium and the search statement based on the image feature vector of the candidate visual medium and the text feature vector of the filtered search statement. This matching degree represents the degree of relevance between the candidate visual medium and the filtered search statement. The greater the matching degree, the more relevant the candidate visual medium is to the filtered search statement. Then, the binary classification model compares the matching degree between the candidate visual medium and the filtered search statement with the classification threshold to determine whether the candidate visual medium matches the search statement, obtaining the Boolean value (bool) 1 corresponding to the candidate visual medium, and thus outputs the Boolean values 1 corresponding to the m candidate visual media. When the matching degree is greater than the classification threshold, it indicates that the candidate visual medium matches the filtered search statement, and the Boolean value 1 corresponding to the candidate visual medium is true. When the matching degree is less than or equal to the classification threshold, it indicates that the candidate visual medium does not match the filtered search statement, and the Boolean value 1 corresponding to the candidate visual medium is false.

[0241] Among them, the binary classification model is a model pre-trained using a training sample set, and it can determine whether these two match based on the image feature vector and the text feature vector. This binary classification model can be trained by the above-mentioned mobile phone or other devices. Below, taking the binary classification model trained by a mobile phone as an example, the training process of the binary classification model will be introduced.

[0242] In some embodiments, when there is a filtered search statement in the above-mentioned preset fine-tuning model vocabulary, it indicates that the filtered search statement corresponds to the branch of the fine-tuning model. Since the branch corresponding to the fine-tuning model does not include the binary classification module, the secondary confirmation module does not need to execute the above S417b.

[0243] When there is no word segmentation of the filtered search statement in the above-mentioned preset fine-tuning model vocabulary, it indicates that the filtered search statement corresponds to the branch of the base model. The text feature vector of the filtered search statement in the above S417b may include the text feature vector of the filtered search statement corresponding to the base model, and the image feature vector of the candidate visual medium in the above S417b may include the image feature vector of the candidate visual medium corresponding to the base model.

[0244] It can be understood that if the input parameters of the secondary confirmation module only include one image feature vector of the candidate visual media and one text feature vector of the filtered search statement, and do not include the image feature vector of the candidate visual media corresponding to the fine-tuning model and the image feature vector of the candidate visual media corresponding to the base model, then the image feature vector of the candidate visual media in the above S417b is the input image feature vector of the candidate visual media, and the text feature vector of the filtered search statement is the input text feature vector of the filtered search statement.

[0245] Exemplarily, the training sample set of the above binary classification model includes multiple pieces of training data, and each piece of training data in the multiple pieces of training data includes a sample image and its corresponding description text. The training sample set can include a positive sample training set (abbreviated as positive samples) and a negative sample training set (abbreviated as negative samples). The sample image in the training data included in the positive samples matches the description text corresponding to the sample image. For example, as Figure 9B shown in the image, the description text corresponding to this image is "a little boy holding a basket". Since Figure 9B the visual content expressed is that the little boy is holding a basket, which matches the corresponding description text, therefore, this image and its corresponding description text can be used as a piece of training data in the positive samples.

[0246] The sample image in the training data included in the negative samples does not match the description text corresponding to the sample image. For example, as Figure 9B shown in the image, the description text corresponding to this image is "flowers". Since Figure 9B the visual content expressed does not match the description text, therefore, this sample image and its corresponding description text can be used as a piece of training data in the negative samples.

[0247] It should be noted that since users generally prefer to use natural language when searching for visual media, the description text corresponding to the sample image also uses natural language, which conforms to the actual search situation.

[0248] To improve the training accuracy, it is necessary to ensure the quantity of the training data in the positive and negative samples. Considering that the efficiency of obtaining positive and negative samples manually is relatively low, the mobile phone can automatically generate positive and negative samples. Below, taking the above binary classification model as a graphic-text matching model as an example, the process of generating positive and negative samples will be continued. As Figure 10 shown, the process is as follows:

[0249] S1. The binary classification module obtains P sample data pairs. Among them, each sample data pair in the P sample data pairs includes a sample image and its corresponding description text.

[0250] S2. The binary classification module extracts sub-texts from the description texts in each sample data.

[0251] Exemplarily, the above-mentioned sub-text may include descriptive phrases and / or nouns in the descriptive text. Exemplarily, the sub-text generally does not include adverbials, verbs, etc. in the descriptive text that do not have corresponding actual visual content. The binary classification module can perform word segmentation on the descriptive text to determine the descriptive phrases and nouns in the descriptive text, and obtain the sub-text in the descriptive text. For example, if the descriptive text is "a little boy sitting in a basket", the sub-texts in this descriptive text are "basket", "little boy", and "boy". Another example, if the descriptive text is "a little boy sitting in a basket and sucking his thumb". After performing word segmentation on this descriptive text, the obtained sub-texts are "basket", "boy", "little boy", and "little boy sucking his thumb", where "little boy sucking his thumb" can represent a descriptive phrase.

[0252] In some embodiments, the binary classification module can use the Han Language Processing Package (HanLP) to perform word segmentation on the descriptive text corresponding to the sample image.

[0253] S3. For each sample image, the binary classification module generates a text set corresponding to the sample image based on the descriptive text corresponding to the sample image and the sub-texts in the descriptive text.

[0254] Among them, the text set corresponding to the sample image represents a description set of the visual content corresponding to the sample image. The text set corresponding to the sample image may include the descriptive text corresponding to the sample image and each sub-text in the descriptive text. For example, if the descriptive text corresponding to the sample image is "a little boy holding a basket", the sub-texts in the descriptive text include "basket", "little boy", and "boy". Correspondingly, the text set corresponding to this image is {"basket", "little boy", "boy", "a little boy holding a basket"}.

[0255] S4. The binary classification module randomly selects a text element from the text sets corresponding to each sample image to obtain the text element 1 corresponding to each sample image.

[0256] Among them, the text elements in the text set can be sub-texts or descriptive texts.

[0257] In some embodiments, the binary classification module can randomly select text elements from the text sets corresponding to each sample image according to a preset ratio. The preset ratio includes the ratio of the selected text elements that are descriptive texts and the ratio of the selected text elements that are sub-texts. For example, if the preset ratio is 8:2, the probability that the selected text element is a descriptive text is 80%, and the probability that the selected text element is a sub-text is 20%. Since generally the image encoder and the text encoder are trained using the image and its corresponding text as a whole, the ratio of the selected text elements that are descriptive texts in the preset ratio will be greater than the ratio of the selected text elements that are sub-texts to ensure the training effect of the model.

[0258] S5. For each sample image, the binary classification module respectively uses the P text elements 1 corresponding to the P sample images as the sample text elements corresponding to the sample image.

[0259] Among them, the number of sample text elements corresponding to one sample image among the P sample images is P. One sample image and one sample text element corresponding to it form a sample, so that multiple samples can be obtained. Since positive and negative sample pairs are required to train the model, after obtaining the samples, it is necessary to distinguish the positive and negative samples in the samples, so as to train the binary classification model with the positive and negative samples. The process of distinguishing the positive and negative samples in the samples will be introduced below.

[0260] In the embodiment of the present application, the binary classification module initially obtains P sample data pairs. These P sample data pairs are all positive samples. However, negative samples are also required for the training of the binary classification model. Therefore, the binary classification module can perform image-text matching on the P sample images and one text element (i.e., text element 1) in the text set corresponding to the P sample images. That is, for each sample image among the P sample images, the binary classification module can use all P text elements 1 as the sample text elements corresponding to the sample image. The sample image and each of the P text elements 1 form a sample, so that P*P samples can be obtained, realizing an increase in the number of samples and thus realizing the rapid generation of samples. To implement the training of the binary classification model, the binary classification model needs to distinguish the positive and negative samples among the P*P samples. The process of distinguishing the positive and negative samples will be continued to be introduced below.

[0261] S6. For each sample text element corresponding to the sample image, the binary classification module determines whether the sample text element belongs to the text set corresponding to the sample image.

[0262] In the embodiment of the present application, for each sample text element corresponding to the sample image, the binary classification module can determine whether the sample text element belongs to the text set corresponding to the sample image, so as to determine whether the sample text element corresponds to the visual content of the sample image, and thus determine whether the sample composed of the sample text element and the sample image is a positive sample.

[0263] In the case where the sample text element is in the text set corresponding to the sample image, it indicates that the sample text element corresponds to (i.e., matches) the visual content of the sample image. That is to say, the sample composed of the sample text element and the sample image is a positive sample. Therefore, the binary classification module can execute S7.

[0264] In the case where the sample text element is not the text set corresponding to the sample image, it indicates that the sample text element does not correspond to (i.e., does not match) the visual content of the image. That is to say, the sample composed of the sample text element and the sample image is not a positive sample. Therefore, the binary classification module can execute S8.

[0265] S7. The binary classification module uses the sample image and the sample text element as positive samples.

[0266] S8. The binary classification module uses the image and the sample text element as negative samples.

[0267] In some embodiments, after obtaining the sample text elements corresponding to each sample image, the binary classification module may generate a corresponding sample matrix. Among them, the i-th row element in the sample matrix represents each sample text element corresponding to the i-th sample image, and the j-th column element in the sample matrix represents the text element 1 corresponding to the j-th sample image, that is, the text element randomly selected from the text set corresponding to the j-th sample image.

[0268] For example, the sample images include image a, image b, image c, and image d. Figure 11A The "1" in the sample matrix shown is the text element 1 corresponding to image a (i.e., the first sample image). And so on, the "6" is the text element 1 corresponding to image d (i.e., the fourth sample image). The first row element 50 (i.e., 1, 2, 5, 6) is the sample text element corresponding to image a. The second column element 51 (i.e., 2, 2, 2, 2) is the text element 1 corresponding to image b.

[0269] Among them, a sample text element in the sample matrix and the sample image corresponding to the row where the sample text element is located form a sample. For example, as described above Figure 11A The "1" in the first row and the first column shown forms a sample with image a.

[0270] In the embodiments of the present application, the binary classification model generates a corresponding sample matrix by mutually matching P sample images and P text elements 1, so that the binary classification module can distinguish whether the sample text elements in the sample matrix belong to positive samples or negative samples, realize the rapid determination of positive and negative samples, improve the determination efficiency of positive and negative samples, and moreover, by generating a label matrix, the comprehensiveness of image-text matching can be ensured, and the situation of missing sample images or text elements 1 can be avoided, such as avoiding the situation that a certain sample image is not matched with a certain text element 1.

[0271] After obtaining the sample matrix, for each sample text element in the sample matrix, the binary classification module needs to determine whether the sample text element matches the sample image corresponding to the row where the sample text element is located. To improve the matching efficiency, the binary classification module can use the text intersection matrix 1 to match with the sample matrix to obtain positive and negative samples. The text intersection matrix 1 represents the intersection of the text sets corresponding to the sample images. Exemplarily, the determination process of the text intersection matrix 1 may include:

[0272] The binary classification module determines the overlapping elements in the text sets corresponding to any two sample images. Then, the binary classification module generates a corresponding text intersection matrix 1. Among them, the s-th element in the t-th row of the text intersection matrix 1 represents the overlapping elements in the text sets between the t-th sample image and the s-th sample image.

[0273] For example, the text set corresponding to image a (as Figure 11B shown) is {1, 2, 3}, the text set corresponding to image b is {2, 3, 4}, the text set corresponding to image c is {5, 6}, and the text set corresponding to image d is {6, 7}. It should be understood that the numbers here are actually corresponding text elements (such as the above sub-texts, description texts).

[0274] Then, the binary classification module can determine the overlapping elements (i.e., the same elements) in the text sets corresponding to any two sample images among image a, image b, image c, and image d. For example, the same elements in the text sets corresponding to image a and image b are 2, 3.

[0275] Then, the binary classification module generates a text intersection matrix 1 as Figure 11C shown. This text intersection matrix 1 includes the overlapping elements in the text sets corresponding to the sample image and each sample image (i.e., image a, image b, image c, and image d). Among them, Figure 11C the element in the first row and first column of the text intersection matrix 1 shown represents the overlapping elements in the text set between image a and image a (i.e., the text set corresponding to image a). The element in the first row and second column represents the overlapping elements in the text set between image a and image b. And so on, Figure 11C the element in the fourth row and fourth column in

[0276] In some embodiments, after the binary classification module obtains the text sets corresponding to each sample image among P sample images, it can calculate the union of the text sets corresponding to each sample image. The union of this text set can include the text elements in all text sets. Then, the binary classification module can assign numbers (such as the above numbers) to each text element in the union of the text sets, so that different text elements in the text set correspond to different numbers, and the same text elements in different text sets correspond to the same number. By assigning numbers to text elements, the efficiency of image-text matching can be improved, and thus the efficiency of generating positive and negative samples can be improved.

[0277] The above introduced the determination process of the text intersection matrix 1. Next, the process of using the text intersection matrix 1 to match with the sample matrix to determine positive and negative samples will be continued.

[0278] The binary classification module takes the intersection of the above-mentioned text intersection matrix 1 and the sample matrix to obtain the text intersection matrix 2.

[0279] After that, for each intersection element in the text intersection matrix 2, the binary classification module determines whether the intersection element is empty.

[0280] When the intersection element is not empty, it indicates that the visual content of the intersection element matches the sample image corresponding to its row. Therefore, the binary classification module can confirm that the intersection element and the sample image corresponding to its row are positive samples.

[0281] When the intersection element is empty, it indicates that the visual content of the intersection element does not match the sample image corresponding to its row. Therefore, the binary classification module can confirm that the element and the sample image corresponding to its row are negative samples, thus realizing the rapid determination of positive and negative samples. Among them, the element corresponding to the sample image can also be called the text corresponding to the sample image.

[0282] Optionally, the binary classification module can use 1 and 0 to distinguish whether the intersection elements in the text intersection matrix 2 belong to positive samples or negative samples, that is, to distinguish whether the sample text elements in the sample matrix belong to positive samples or negative samples. When the intersection element in the text intersection matrix 2 is empty, the binary classification module can set the label corresponding to the intersection element to 0, that is, label the sample text element corresponding to the intersection element with 0.

[0283] When the intersection element in the text intersection matrix 2 is not empty, the binary classification module can set the label corresponding to the element to 1, that is, label the sample text element corresponding to the intersection element with 1, thereby obtaining the positive and negative sample label matrix and realizing the setting of labels for the sample text elements in the sample matrix. The position of the intersection element in the text intersection matrix 2 is the same as the position of the sample text element corresponding to the intersection element in the sample matrix.

[0284] Based on this, the binary classification module can use the sample text element corresponding to the label 0 and the image corresponding to the row where the sample text element is located as a piece of training data in the negative samples, and use the sample text element corresponding to the label 1 and the image corresponding to the row where the sample text element is located as a piece of training data in the positive samples, realizing the batch determination of multiple positive and negative sample training data.

[0285] For example, as Figure 11D shown, the text intersection matrix 1 and the sample matrix are intersected to obtain the text intersection matrix 2. After that, the binary classification module can determine whether the intersection elements in the text intersection matrix 2 are empty, thereby determining the positive and negative sample label matrix corresponding to the sample matrix. As Figure 11DThe label in the first row and first column of the positive and negative sample label matrix is 1. The element in the sample matrix corresponding to this label is the sample text element "1" in the first row and first column. This "1" and image a form a positive sample. As Figure 11D The label in the second row and first column of the positive and negative sample label matrix is 0. The sample text element in the second row and first column of the sample matrix corresponding to this label is "1". This "1" and image b form a negative sample.

[0286] It should be noted that generally, the elements on the diagonal of the sample matrix and the corresponding sample images in their rows form positive samples.

[0287] In the embodiments of the present application, the binary classification module obtains the text intersection matrix 2 by intersecting the text intersection matrix 1 and the sample matrix, so as to quickly determine positive and negative samples by using whether the elements in the text intersection matrix 2 are empty, improve the determination efficiency of positive and negative samples, and ensure the comprehensiveness and integrity of the distinction between positive and negative samples, so that the model can be trained quickly.

[0288] The process of determining positive and negative samples is introduced above. Next, the process of training a binary classification model using positive and negative samples will be continued.

[0289] S9. The binary classification module trains the text-image matching model using positive and negative samples to obtain a trained text-image matching model.

[0290] Exemplarily, the text-image matching model can be a multilayer perceptron (MLP) model. As Figure 11E shown, the binary classification module processes the sample text elements in positive and negative samples using a text encoder to obtain text feature vectors of the sample text elements. And, the binary classification module processes the sample images in positive and negative samples using an image encoder to obtain image feature vectors of the sample images. Then, for each sample in positive and negative samples, the binary classification module can concatenate the text feature vector of the sample text element in the sample and the image feature vector of the sample image in the sample. Then, the binary classification module can input the concatenated text feature vector and image feature vector into the MLP model to train the MLP model. During the training process, the binary classification module can use a loss function to test the difference between the predicted value output by the trained MLP model and the actual value. The predicted value indicates whether the predicted sample image matches the text. The actual value indicates the actual matching situation between the sample image and the text.

[0291] In some embodiments, the above loss function can be a focal loss function. Specifically, the focal loss function can be FL(p t )=-at 1(1 - p t ) γ log(p t ). Among them, \(p_t\) represents the probability of the matching between the predicted sample image and its corresponding text by the image - text matching model, that is, the probability that the predicted sample image and its corresponding text are positive samples. \(a_{t1}\) is a factor for adjusting the weights of positive and negative samples, which can be set according to the number of positive and negative samples to adjust the imbalance between the number of positive and negative samples. For example, \(a_{t1}\) is 0.1. \(\gamma\) is a adjustment factor used to reduce the loss contribution of easily distinguishable samples. It should be understood that the larger \(p_t\) is, the smaller the value of the loss function is.

[0292] Among them, optionally, the above - mentioned image encoder can be a stacked auto - encoder, and the text encoder can be a count vectorizer.

[0293] It should be noted that the above - introduced focal loss function is only an example of the loss function, and this loss function can also be other types of loss functions. For example, this loss function is a loss function of the BCE class. In addition, the above - mentioned MLP model is only an example of the image - text matching model, and this image - text matching model can also be other deep learning models, and the present application does not limit it.

[0294] In some embodiments, the above - mentioned positive and negative sample label matrix can be used when calculating the value of the loss function.

[0295] In some embodiments, the use of the above - mentioned image - text matching model to perform secondary confirmation on candidate visual media is only an example, and this image - text matching model can also be directly used for searching visual media. For example, after the user inputs a search statement, the image - text matching model can directly use the text feature vector of this search statement (or filtered search statement) and the image feature vectors of each visual media on the mobile phone to determine whether the visual media matches the search statement, and obtain the corresponding boolean value.

[0296] The process of using the binary classification module to determine whether the candidate visual media matches the filtered search statement is introduced above. Next, in combination with S417c, the process of using the dynamic threshold module to determine whether the candidate visual media matches the filtered search statement will be introduced.

[0297] S417c: The binary classification module inputs the filtered search statement and the similarity between each candidate visual media and the filtered search statement into the dynamic threshold module, and obtains the boolean value 2 corresponding to each candidate visual media output by the dynamic threshold module.

[0298] In the embodiments of the present application, generally, the similarity threshold 1 corresponding to search statements of different lengths is different. The longer the length of the search statement, the higher the similarity degree between the visual content of the visual media and the search statement needs to be, and correspondingly, the higher the similarity threshold 1 needs to be. Therefore, the secondary confirmation module can use the dynamic threshold module to determine the dynamic threshold that matches the length of the filtered search statement, that is, determine the similarity threshold 1 corresponding to the filtered search statement, so as to determine the Boolean value 2 corresponding to each candidate visual media by using the similarity threshold 1 corresponding to the filtered search statement.

[0299] Exemplarily, as Figure 12 shown, the process for the dynamic threshold module to determine the Boolean value 2 corresponding to each candidate visual media may include: First, the dynamic threshold module can determine the length of the filtered search statement through the filtered search statement. After that, based on the length of the filtered search statement, the dynamic threshold model combines t = parameter 1 * L + parameter 2 to determine the similarity threshold 1 corresponding to the filtered search statement. Among them, the above parameters 1 and 2 are preset parameters. For example, parameter 1 is 0.05 and parameter 2 is 0.33. Optionally, the dynamic threshold module can also determine the length of the filtered search statement through the word segmentation result of the filtered search statement.

[0300] After that, for each candidate visual media among the m candidate visual media, the dynamic threshold module can compare the size between the similarity between the candidate visual media and the filtered search statement and the similarity threshold 1 corresponding to the filtered search statement. In the case where the similarity between the candidate visual media and the filtered search statement is less than the similarity threshold 1, it indicates that the similarity degree between the visual content of the candidate visual media and the filtered search statement is relatively low, and the dynamic threshold module can determine that the Boolean value 2 corresponding to the candidate visual media is false.

[0301] In the case where the similarity between the candidate visual media and the filtered search statement is greater than or equal to the similarity threshold 1, it indicates that the similarity degree between the visual content of the candidate visual media and the filtered search statement is relatively high, and the dynamic threshold module can determine that the Boolean value 2 corresponding to the candidate visual media is true.

[0302] Among them, optionally, as shown above Figure 12 shown, the value range of t can be greater than or equal to parameter 3, that is, min(t, parameter 3), and the value range of t can be less than or equal to parameter 4. That is, max(t, parameter 4). Among them, parameter 3 and parameter 4 are preset.

[0303] In some embodiments, since the dynamic threshold module corresponding to S417c belongs to the branch corresponding to the base model, the similarity between the candidate visual media in S417c and the filtered search statement includes the similarity between the candidate visual media corresponding to the base model and the filtered search statement (i.e., similarity 1). Accordingly, the dynamic threshold module can determine whether the similarity 1 between the candidate visual media and the filtered search statement is less than the similarity threshold 1. When the similarity 1 is greater than or equal to the similarity threshold 1, the dynamic threshold module can determine that the Boolean value 2 corresponding to the candidate visual media is true. In the case where the similarity 1 is less than the similarity threshold 1, the dynamic threshold module can determine that the Boolean value 2 corresponding to the candidate visual media is false.

[0304] It can be understood that if the input parameter of the secondary confirmation module only includes one similarity between the candidate visual media and the filtered search statement, and does not include the similarity between the candidate visual media corresponding to the fine-tuning model and the filtered search statement and the similarity between the candidate visual media corresponding to the base model and the filtered search statement, then the image feature vector of the candidate visual media in S417c is the similarity between the input candidate visual media and the filtered search statement.

[0305] In addition, the similarity between the candidate visual media and the filtered search statement can also be determined by other multimodal models, and the present application does not limit it.

[0306] The process of determining the Boolean value corresponding to the candidate visual media using the dynamic threshold module is introduced above. Next, the process of determining whether the candidate visual media matches the filtered search statement using the label confirmation module will be continued with reference to S417d.

[0307] S417d. The secondary confirmation module inputs the tokenization result of the above filtered search statement, label 1 included in the filtered search statement, and label 2 of each candidate visual media into the label confirmation module, and obtains the Boolean value 3 corresponding to each candidate visual media output by the label confirmation module.

[0308] In the embodiments of the present application, the secondary confirmation module can use the label confirmation module to determine whether label 2 of the candidate visual media exists in the tokenization of the filtered search statement or label 1 included in the filtered search statement, so as to determine whether the candidate visual media matches the filtered search statement, and thus determine the Boolean value 3 corresponding to the candidate visual media. Exemplarily, as Figure 13 shown, the process of the label confirmation model determining the Boolean value 3 corresponding to the candidate visual media can include:

[0309] First, for each of the m candidate visual media, the label confirmation module can determine whether the label 2 of the candidate visual media includes the word segments in the word segmentation result of the filtered search statement, and determine whether the label 2 of the candidate visual media includes the label 1 included in the filtered search statement, that is, determine whether the candidate visual media includes the visual content corresponding to the filtered search statement.

[0310] In the case where the label 2 of the above candidate visual media includes at least one word segment of the filtered search statement, or the label 2 of the candidate visual media includes at least one label 1 included in the filtered search statement, it indicates that the candidate visual media hits the filtered search statement. The label confirmation module can determine that the candidate visual media may be the visual media required by the user. Therefore, the label confirmation module can determine the Boolean value 3 corresponding to the candidate visual media as true. For example, the label 1 of the filtered search statement includes "child" and "flower", and the word segmentation result of the filtered search statement includes the word segments "child", "holding", and "flower". Then, in the case where the label 2 of the candidate visual media includes "child", "holding" or "flower", or the label 2 of the candidate visual media includes "child" or "flower", it is determined that the Boolean value 3 corresponding to the candidate visual media is true.

[0311] In the case where the label 2 of the above candidate visual media does not include all the word segments of the filtered search statement, and the label 2 of the candidate visual media does not include all the label 1 included in the filtered search statement, it indicates that the candidate visual media does not hit the filtered search statement, and the candidate visual media may not be the visual media required by the user. Therefore, the label confirmation module can determine the Boolean value 3 corresponding to the candidate visual media as false, so as to obtain the Boolean value 3 of the m candidate visual media.

[0312] Optionally, in order to improve the accuracy of the search results, the label confirmation module can determine that the Boolean value 3 corresponding to the candidate visual media is true when the label 2 of the candidate visual media includes all the word segments of the filtered search statement, or includes all the label 1 included in the filtered search statement.

[0313] When the label 2 of the candidate visual media does not include at least one word segment of the filtered search statement and does not include at least one label 1 included in the filtered search statement, it is determined that the Boolean value 3 corresponding to the candidate visual media is false.

[0314] The process of the secondary confirmation module using the label confirmation module to determine the Boolean value 3 corresponding to each candidate visual media is introduced above. Next, the process of the secondary determination module using the whitelist threshold module to determine whether the candidate visual media matches the filtered search statement will be continued.

[0315] The secondary confirmation module inputs the text feature vector of the filtered search statement, the candidate visual media, and the similarity between the candidate visual media and the filtered search statement into the whitelist threshold module, and obtains the Boolean value 4 corresponding to each candidate visual media output by the whitelist threshold module.

[0316] In the embodiment of the present application, the secondary confirmation module can determine whether the filtered search statement hits the dictionary through the whitelist threshold module, so as to determine whether there is a similarity threshold 2 corresponding to the filtered search statement, and accurately determine the similarity threshold. Exemplarily, the whitelist threshold module can determine whether the dictionary contains the filtered search statement. If it exists, the whitelist threshold module takes effect, and the whitelist threshold module can use the similarity threshold corresponding to the filtered search statement in the dictionary as the similarity threshold 2 corresponding to the filtered search statement, so as to screen the candidate visual media matching the filtered search statement by using the similarity threshold 2. If it does not exist, the whitelist threshold module does not take effect, that is, there is no need to use the whitelist threshold module to determine the Boolean value 4 corresponding to each candidate visual media. Among them, the dictionary includes a key and the value corresponding to the key. The key represents a preset search text, and the value is the similarity threshold corresponding to the preset search text. The key in the dictionary can be the search text frequently input by the user, that is, the search statement that has been pre-tested.

[0317] In some embodiments, there are corresponding dictionaries for the branches corresponding to the fine-tuning model and the base model. As Figure 14A shown, when the filtered search statement is the branch corresponding to the base model, the whitelist threshold module can determine whether the dictionary corresponding to the base model includes the filtered search statement, that is, determine whether the dictionary corresponding to the base model has the same key as the filtered search statement.

[0318] If the dictionary corresponding to the base model includes the filtered search statement, it indicates that the filtered search statement hits the dictionary corresponding to the base model, that is, the dictionary corresponding to the base model has the same key as the filtered search statement. Then, the whitelist threshold module can use the value corresponding to the filtered search statement in the dictionary corresponding to the base model as the similarity threshold 2.

[0319] After that, for each of the m candidate visual media, the whitelist threshold module can compare the similarity between the candidate visual media corresponding to the base model and the filtered search threshold with the similarity threshold 2. When the similarity between the candidate visual media corresponding to the base model and the filtered search threshold is less than the similarity threshold 2, the whitelist threshold module can determine that the Boolean value 4 corresponding to the candidate visual media is false. When the similarity is greater than or equal to the similarity threshold 2, it indicates that the visual content of the candidate visual media is highly similar to the filtered search statement, and the whitelist threshold module can determine that the Boolean value corresponding to the candidate visual media is true.

[0320] Correspondingly, when filtering the search statement for the branch corresponding to the base model, the above S418 can be that for each candidate visual medium, when the Boolean value corresponding to the candidate visual medium has a true value, the secondary confirmation module (such as the fusion module in the secondary confirmation module) can use the candidate visual medium as Visual Medium 1. Among them, the Boolean value corresponding to the candidate visual medium includes the above Boolean Value 1, Boolean Value 2, Boolean Value 3, and Boolean Value 4. That is to say, the secondary confirmation module can take the union of the Boolean Value 1, Boolean Value 2, Boolean Value 3, and Boolean Value 4 corresponding to the candidate visual medium to obtain the Boolean value corresponding to the candidate visual medium.

[0321] When the Boolean value corresponding to the candidate visual medium is all false, it indicates that the Boolean Value 1, Boolean Value 2, Boolean Value 3, and Boolean Value 4 corresponding to the candidate visual medium are all false, and the secondary confirmation module (such as the fusion module in the secondary confirmation module) can not use the candidate visual medium as Visual Medium 1.

[0322] It should be noted that since both the above whitelist threshold module and the above dynamic threshold module are used to confirm the similarity threshold corresponding to the filtered search statement, and the similarity threshold determined by the whitelist threshold module is more accurate. Therefore, when the above whitelist threshold module takes effect, that is, when using the similarity threshold 2 corresponding to the filtered search statement to determine the Boolean value corresponding to the candidate visual medium, the dynamic threshold module can be ineffective, and the secondary confirmation module can not use the similarity threshold 1 corresponding to the filtered search statement to determine the Boolean value corresponding to the candidate visual medium, avoiding unnecessary determination of the similarity threshold, and thus avoiding unnecessary screening of candidate visual media.

[0323] In some embodiments, the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module in the above secondary confirmation module can determine the Boolean value corresponding to the candidate visual medium in parallel or serially. Whether it is determined in parallel or serially, when one module in the secondary determination module determines that the Boolean value corresponding to the candidate visual medium is true, other modules do not need to continue to determine the Boolean value corresponding to the candidate visual medium, that is, do not need to input the relevant information of the candidate visual medium to other models, avoiding unnecessary determination of the Boolean value and ensuring the transmission cost.

[0324] In addition, the above-mentioned secondary confirmation module including the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module is only an example. The secondary confirmation module may include one or more of the binary classification module, the dynamic threshold module, the label confirmation model, and the whitelist threshold module to improve the search efficiency of visual media. For example, if the secondary confirmation module includes one of the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module, such as the secondary confirmation module includes the binary classification module, correspondingly, the Boolean value corresponding to the above-mentioned candidate visual media may include Boolean value 1. For another example, if the secondary confirmation module includes two of the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module, such as the secondary confirmation module includes the binary classification module and the dynamic threshold module. Correspondingly, the Boolean values corresponding to the above-mentioned candidate visual media may include Boolean value 1 and Boolean value 2. For another example, if the secondary confirmation module includes three of the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module, such as the secondary confirmation module includes the binary classification module, the dynamic threshold module, and the label confirmation module. Correspondingly, the Boolean values corresponding to the above-mentioned candidate visual media may include Boolean value 1, Boolean value 2, and Boolean value 3.

[0325] Moreover, the content included in the image information of the above-mentioned candidate visual media and the information of the filtered search statement, that is, the input parameters of the secondary confirmation module are only an example, and it can be adaptively set according to the modules included in the secondary confirmation module. For example, if the secondary confirmation module includes a binary classification model, the image information of the above-mentioned candidate visual media may include the image feature vector of the candidate visual media, and the information of the filtered search statement may include the text feature vector of the filtered search statement.

[0326] The above introduced the process in which the secondary confirmation module can sequentially use the binary classification module, the dynamic threshold module, the label confirmation module, and the whitelist threshold module to determine the Boolean value corresponding to the candidate visual media when the filtered search statement corresponds to the branch of the base model. Next, the process in which the secondary confirmation module can sequentially use the label confirmation module and the whitelist threshold module to determine the Boolean value corresponding to the candidate visual media when the filtered search statement corresponds to the branch of the fine-tuning model will be continued.

[0327] S417f. The secondary confirmation module inputs the word segmentation result of the above-mentioned filtered search statement, label 1 included in the filtered search statement, and label 2 of each candidate visual media into the label confirmation module, and obtains the Boolean value 5 corresponding to each candidate visual media output by the label confirmation module.

[0328] Among them, the implementation process of S417f can refer to the implementation process of the above-mentioned S417d, and will not be elaborated here.

[0329] S417g. The secondary confirmation module inputs the similarity between the filtered search statement and the candidate visual media into the whitelist threshold module, and obtains the boolean value 4 corresponding to each candidate visual media output by the whitelist threshold module.

[0330] Among them, the implementation process of S417g can refer to the implementation process of the above S417e. As Figure 14B shown, when the filtered search statement is the branch corresponding to the fine-tuning model, the whitelist threshold module can determine whether the dictionary corresponding to the fine-tuning model includes the filtered search statement, that is, determine whether there is a key in the dictionary corresponding to the fine-tuning model that is the same as the filtered search statement, so as to determine whether the whitelist threshold module takes effect.

[0331] After the whitelist threshold module takes effect, the secondary confirmation module can compare the similarity between the candidate visual media corresponding to the fine-tuning model and the filtered search statement with the value corresponding to the filtered search statement in the dictionary corresponding to the fine-tuning model, so as to determine the boolean value corresponding to the candidate visual media.

[0332] Correspondingly, when the filtered search statement is the branch corresponding to the fine-tuning model, the above S418 can be that for each candidate visual media, when the boolean value corresponding to the candidate visual media is true, the secondary confirmation module (such as the fusion module in the secondary confirmation module) can use the candidate visual media as the visual media 1. Among them, the boolean value corresponding to the candidate visual media includes the above boolean value 5 and the above boolean value 6, that is to say, the secondary confirmation module can take the union of the boolean value 5 and the boolean value 6 corresponding to the candidate visual media to obtain the boolean value corresponding to the candidate visual media.

[0333] In the embodiments of the present application, after obtaining the candidate visual media, the mobile phone uses the secondary confirmation module to continue to determine the boolean value corresponding to the candidate visual media, so as to determine whether the candidate visual media matches the filtered search statement, so as to select the candidate visual media that matches the filtered search statement from the candidate visual media, obtain the corresponding search result, ensure the accuracy of the search result, and thus ensure the user satisfaction.

[0334] It should be noted that the operations performed by the above modules or models are only examples, and the operations performed by the above modules can also be performed by other modules in the mobile phone. The present application does not limit it. In addition, the operations actually performed by the above modules or models are performed by the mobile phone.

[0335] In some embodiments, when the mobile phone determines the candidate visual media or performs secondary confirmation on the candidate visual media, it may not filter the non-visual semantic entities in the search statement first, but directly use the text feature vector of the search statement to determine the candidate visual media, or perform secondary confirmation on the candidate media.

[0336] In some embodiments, the search and storage of the above visual media are carried out under the authorization of the user, including but not limited to, before the user uses this function, notifying and reminding the user to read the relevant user agreement (notification), and signing the agreement (authorization) including authorizing the relevant user information.

[0337] In some embodiments, the present application provides a computer storage medium including computer instructions, which, when running on an electronic device, cause the electronic device to execute the method as described above.

[0338] In some embodiments, the present application provides a computer program product, which, when running on an electronic device, causes the electronic device to execute the method as described above.

[0339] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions as described in the embodiments of the present application are generated in whole or in part.

[0340] It should be understood that the "embodiments" mentioned throughout the specification mean that the specific features, structures, or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, the various embodiments throughout the specification do not necessarily refer to the same embodiments. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0341] It should also be understood that in the present application, "when...", "if", and "in case" all mean that under certain objective circumstances, the UE or the base station will perform corresponding processing, which does not limit the time, and does not require the UE or the base station to have a judgment action when implemented, nor does it mean that there are other limitations.

[0342] Those of ordinary skill in the art can understand that the various numerical numbers such as the first and second involved in the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application, nor do they represent the order of precedence.

[0343] In this application, elements expressed in the singular are intended to mean "one or more", rather than "one and only one", unless otherwise specified. In this application, unless otherwise specified, "at least one" is intended to mean "one or more", and "a plurality of" is intended to mean "two or more".

[0344] The term "and / or" in this document is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A can be singular or plural, and B can be singular or plural.

[0345] The term "at least one of..." or "at least one kind of..." in this document means all or any combination of the items listed. For example, "at least one of A, B, and C" can represent: A exists alone, B exists alone, C exists alone, A and B exist simultaneously, B and C exist simultaneously, and A, B, and C exist simultaneously. Here, A can be singular or plural, B can be singular or plural, and C can be singular or plural.

[0346] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0347] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0348] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0349] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0350] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0351] If the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0352] The same or similar parts among the various embodiments of this application can be referred to each other. In each embodiment of this application, as well as in each implementation manner / implementation method / realization method in each embodiment, if there is no special description and logical conflict, the terms and / or descriptions among different embodiments, as well as among the various implementation manners / implementation methods / realization methods in each embodiment, are consistent and can be referenced to each other. The technical features in different embodiments, as well as in the various implementation manners / implementation methods / realization methods in each embodiment, can be combined according to their internal logical relationships to form new embodiments, implementation manners, implementation methods, or realization methods. The implementation manners of this application described above do not constitute a limitation on the protection scope of this application.

[0353] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims. In short, the above description is only the preferred embodiment of the technical solution of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A visual media search method applicable to an electronic device, characterized in that, the method includes: Obtain P sample data pairs, where each of the P sample data pairs includes a sample image and a corresponding description text; For each sample image, select a first text element from the text set corresponding to the sample image to obtain the first text element corresponding to the sample image; wherein, the text set corresponding to the sample image includes the description text corresponding to the sample image and a first sub-text in the description text; Perform image-text matching on the P sample images and the first text elements corresponding to the P sample images to obtain a P*P sample matrix; Train an image-text matching model based on the sample matrix and the sample images; Display a first interface; the first interface includes a search box; Receive a search statement input in the search box; the search statement includes one or more sub-texts; Display search results, where the search results correspond to a first visual media; the first visual media represents a visual media whose visual content determined by the image-text matching model matches the search statement.

2. The method according to claim 1, characterized in that, the performing image-text matching on the P sample images and the first text elements corresponding to the P sample images to obtain a P*P sample matrix includes: For each sample image, use the first text elements corresponding to the P sample images as the sample text elements corresponding to the sample image; Generate the sample matrix based on the sample text elements corresponding to each sample image; wherein, the elements in the i-th row of the sample matrix represent the respective sample text elements corresponding to the i-th sample image; any sample text element in the sample matrix and the sample image corresponding to the row where the sample text element is located form a sample.

3. The method according to claim 1 or 2, characterized in that, the training an image-text matching model based on the sample matrix and the sample images includes: Determine positive and negative samples based on the sample matrix, the text set corresponding to the sample image, and the sample image; Train an image-text matching model based on the positive and negative samples.

4. The method according to claim 3, characterized in that, the determining positive and negative samples based on the sample matrix, the text set corresponding to the sample image, and the sample image includes: For each sample text element in the sample matrix, when the sample text element belongs to the text set corresponding to the sample image corresponding to the row where the sample text element is located, determine that the sample text element and the sample image corresponding to the row where the sample text element is located belong to positive samples; When the sample text element does not belong to the text set corresponding to the sample image corresponding to the row where the sample text element is located, determine that the sample text element and the sample image corresponding to the row where the sample text element is located belong to negative samples.

5. The method according to claim 3, characterized in that, the determining positive and negative samples based on the sample matrix, the text set corresponding to the sample image, and the sample image includes: Intersect the first text intersection matrix and the sample matrix to obtain a second text intersection matrix; the first text intersection matrix is determined based on the same elements in the text sets corresponding to any two of the sample images; When the s-th intersection element in the t-th row of the second text intersection matrix is empty, determine that the s-th intersection element in the t-th row of the sample matrix and the t-th sample image belong to negative samples; When the s-th intersection element in the t-th row of the second text intersection matrix is not empty, determine that the s-th intersection element in the t-th row of the sample matrix and the t-th sample image belong to positive samples.

6. The method according to any one of claims 1 to 5, wherein, Before performing the graphic-text matching on P sample images and the first text elements corresponding to the P sample images to obtain a P*P sample matrix, the method further includes: Determine the union of the text sets corresponding to each sample image; Assign numbers to each text element in the union of the text sets; The performing the graphic-text matching on P sample images and the first text elements corresponding to the P sample images to obtain a P*P sample matrix includes: Performing graphic-text matching on the numbers corresponding to the P sample images and the first text elements corresponding to the P sample images to obtain a P*P sample matrix.

7. The method according to any one of claims 1 to 6, wherein, For each sample image, the selecting the first text element from the text set corresponding to the sample image includes: Randomly selecting text elements from the text set corresponding to each sample image according to a preset ratio; wherein, the preset ratio includes the ratio of the selected first text elements that are descriptive texts and the ratio of the selected first text elements that are first sub-texts.

8. The method according to any one of claims 1 to 7, wherein, The first visual media represents candidate visual media determined by a graphic-text matching model, and the visual content of which matches the search statement; the candidate visual media represents visual media with a similarity greater than a first threshold to the search statement.

9. The method according to any one of claims 1 to 6, wherein, The first visual media represents visual media determined by a graphic-text matching model, and the visual content of which matches the filtered search statement; the filtered search statement is obtained by filtering non-visual semantic entities in the search statement, and the non-visual semantic entities do not correspond to visual content.

10. An electronic device, wherein, The electronic device includes a display screen, a memory, and one or more processors; the display screen, the memory, and the processor are coupled; the display screen is used to display images generated by the processor, the memory is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the electronic device is caused to execute the method according to any one of claims 1 to 9.

11. A computer storage medium, wherein, Comprising computer instructions that, when executed on an electronic device, cause the electronic device to perform the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Training method, image recognition method and device, equipment and readable storage medium

    CN116092101A

  • Image-text retrieval model training method, image-text retrieval method, image-text retrieval device and image-text retrieval equipment

    CN116226353A

  • Data processing method, image-text retrieval method, image classification method and related equipment

    CN116226688A

  • Visual language understanding task processing method and system

    CN116432026A

  • Cross-modal retrieval method based on multilevel fine-grained semantic alignment

    CN116775929A