Visual media searching method and device and storage medium

By using natural picture understanding model and semantic subject technology in electronic devices and adjusting the sorting and filtering of visual media searches, the problem of inaccurate visual media searches in the prior art is solved, and higher search accuracy and user experience are achieved.

CN120067370APending Publication Date: 2025-05-30HONOR DEVICE CO LTD

Patent Information

Application Number
CN202311565684.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to accurately recall pictures or videos related to users' complex search statements in visual media searches, resulting in inaccurate search results.

Method used

By implementing a visual media search method in an electronic device, the natural picture understanding model is used to obtain the visual semantic vectors of the visual media, and combined with the semantic subject in the search statement, the sorting and filtering of the recalled visual media files are adjusted to improve the accuracy of the search results.

Benefits of technology

It improves the accuracy and user experience of visual media searches, and can recall visual media files that meet users' complex search statements more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067370A_ABST
    Figure CN120067370A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a visual media searching method and device and a storage medium. The method is suitable for the electronic equipment and comprises the following steps: displaying a first interface; the first interface comprises a search box; receiving a search operation of a search statement input to the search box; according to a first sentence semantic vector of the search statement, determining M candidate visual media files of which the visual semantic vectors are matched with the first sentence semantic vector from a plurality of visual media; the visual semantic vector of each visual media file is obtained by carrying out semantic understanding on an image or an image frame of the visual media file by utilizing a natural picture understanding model; determining N to-be-matched dimensions according to a semantic subject related to the visual content in the search statement; determining a search result of the search statement according to the matching degree of the visual contents of the M candidate visual media files on the N to-be-matched dimensions; and displaying the search result. According to the technical scheme provided by the embodiment of the invention, the search accuracy can be improved, and the user search experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of terminals, and in particular, to a visual media search method, device, and storage medium. Background Art

[0002] With the popularization of smart terminals, more and more users use smart terminals, such as mobile phones, to take pictures and videos, and store the taken pictures and videos in the picture gallery of the electronic device, so as to record every bit of life. In addition, users can also download pictures, take screenshots of the mobile phone interface, and store the downloaded pictures and screenshots in the picture gallery of the electronic device.

[0003] In order to facilitate users to manage and view pictures in the terminal, a picture management function and a picture search function are configured in the picture gallery application or other similar applications of the terminal device. For example: the picture gallery application in the terminal can classify the pictures in the terminal according to information such as the shooting time and location of the pictures to generate corresponding albums, and users can view relevant pictures by searching for information such as time and location. Summary of the Invention

[0004] Multiple aspects of this application provide a visual media search method, device, and storage medium, which can improve the search accuracy and the user search experience.

[0005] In a first aspect, a visual media search method applicable to an electronic device is provided, including:

[0006] Display a first interface; the first interface includes a search box;

[0007] Receive a search operation on a search statement input into the search box;

[0008] According to the first sentence semantic vector of the search statement, determine M candidate visual media files whose visual semantic vectors match the first sentence semantic vector from multiple visual media; the visual semantic vector of each visual media file is obtained by performing semantic understanding on the image or image frame of the visual media file using a natural picture understanding model; where M>1 and is an integer;

[0009] According to the semantic entity related to visual content in the search statement, determine N dimensions to be matched; N≥1 and is an integer;

[0010] According to the matching degree of the visual content of the M candidate visual media files on the N dimensions to be matched, determine the search result of the search statement;

[0011] Display the search result.

[0012] The above-mentioned M candidate visual media files are obtained by matching the visual semantic vectors of the visual media files with the sentence semantic vectors of the search statement. That is to say, when recalling the above-mentioned M candidate visual media files, the matching degree between the visual semantics of the visual media files and the entire search statement of the user is considered, rather than the matching degree between the visual semantics of the visual media files and the semantic entities that the user visually focuses on in the search statement, resulting in the recall of some inaccurate photos and videos. Therefore, in this solution, the above-mentioned M candidate visual media files are fine-tuned by means of the matching degree between the visual semantics of the visual media and the "semantic entities related to visual content" to obtain more accurate search results. Among them, the fine-tuning may include: reordering or filtering.

[0013] In a possible implementation manner, the method further includes:

[0014] Match the search terms in the search statement with a preset plurality of tags to determine whether the search terms belong to the semantic entities related to the tags; the tags are used to describe visual content;

[0015] Determine the semantic entities related to visual content in the search statement as the semantic entities related to visual content in the search statement.

[0016] That is, the search terms that match a certain tag among the preset plurality of tags belong to the semantic entities related to the tag. Since the tag is related to visual content, the semantic entities related to the tag also belong to the semantic entities related to visual content. In this solution, by presetting a plurality of tags, the semantic entities that the user visually focuses on can be relatively simply determined from the search statement.

[0017] In a possible implementation manner, according to the semantic entities related to visual content in the search statement, N to-be-matched dimensions are determined, including:

[0018] Use the N semantic entities related to visual content in the search statement as N to-be-matched dimensions; or

[0019] Use (N - 1) semantic entities related to visual content in the search statement and the search statement as N to-be-matched dimensions.

[0020] In a possible implementation manner, according to the matching degrees of the visual content of the M candidate visual media files on the N to-be-matched dimensions, the search result of the search statement is determined, including:

[0021] Obtain the weights of the N to-be-matched dimensions; among them, the weight of the jth to-be-matched dimension is positively correlated with the variation degree of the matching degree of the visual content of the M candidate visual media files on the jth to-be-matched dimension; j is an integer, and the value of j ranges from 1 to N in sequence;

[0022] According to the weights of the N dimensions to be matched, the matching degrees of the visual content of the i-th candidate visual media file on each dimension to be matched among the N dimensions to be matched are weighted and summed to obtain the comprehensive matching degree of the i-th candidate visual media file; i is an integer, and the values of i range from 1 to M in sequence;

[0023] According to the comprehensive matching degrees of the M candidate visual media files, determine the search result of the search statement.

[0024] In this solution, the roles played by different dimensions to be matched in fine-tuning are also different, that is, the weights of different dimensions to be matched are different. The weight of each dimension to be matched is determined according to the degree of variation of the matching degrees of the visual content of the M candidate visual media files on this dimension to be matched. The greater the degree of variation, the greater the role played by the dimension to be matched in fine-tuning. Therefore, its weight is greater. In this way, the effect of fine-tuning can be improved to improve the search accuracy.

[0025] In a possible implementation manner, the method further includes:

[0026] According to the information entropy of the matching degrees of the visual content of the M candidate visual media files on the j-th dimension to be matched, determine the degree of variation of the matching degrees of the visual content of the M candidate visual media files on the j-th dimension to be matched; the degree of variation is negatively correlated with the information entropy;

[0027] According to the degree of variation of the matching degrees of the visual content of the M candidate visual media files on the j-th dimension to be matched, determine the weight of the j-th dimension to be matched.

[0028] In this solution, the information entropy is used to characterize the degree of variation.

[0029] In a possible implementation manner, the method further includes:

[0030] According to the vector similarity between the visual semantic vectors of the M candidate visual media files and the semantic vectors to be matched of the j-th dimension to be matched, determine the initial matching degrees of the M candidate visual media on the j-th dimension to be matched;

[0031] Perform normalization processing on the initial matching degrees of the M candidate visual media on the j-th dimension to be matched to obtain the matching degrees of the M candidate visual media on the j-th dimension to be matched.

[0032] In this solution, through the normalization operation, it can be ensured that the matching degrees of the visual content of the M candidate visual media files on different dimensions to be matched have a unified dimension.

[0033] In a possible implementation, determining a search result of the search statement according to a comprehensive matching degree of the M candidate visual media files includes:

[0034] Determining a plurality of first visual media files whose comprehensive matching degrees meet a preset requirement from the M candidate visual media files;

[0035] Determining a search result of the search statement according to the plurality of first visual media files.

[0036] In a possible implementation, determining a search result of the search statement according to a plurality of first visual media files includes:

[0037] If the search statement includes a semantic entity related to time, filtering the plurality of first visual media files according to the semantic entity related to time and a collection time attribute of the plurality of first visual media files;

[0038] If the search statement includes a semantic entity related to a location, filtering the plurality of first visual media files according to the semantic entity related to the location and a collection location attribute of the plurality of first visual media files;

[0039] If the search statement includes a semantic entity related to a person relationship, filtering the plurality of first visual media files according to the semantic entity related to the person relationship and a person relationship attribute of the plurality of first visual media files; and / or

[0040] If the search statement includes a semantic entity related to a person name, filtering the plurality of first visual media files according to the semantic entity related to the person name and a person name attribute of the plurality of first visual media files.

[0041] In this solution, time filtering, location filtering, person relationship filtering, and person name filtering are performed on the recalled visual media.

[0042] In a second aspect, the present application provides an electronic device, including: a memory, a processor, and a display, where

[0043] The memory is used to store a program;

[0044] The processor is coupled to the memory and the display, and is used to execute the program stored in the memory to implement any one of the methods.

[0045] In a third aspect, the present application provides a computer-readable storage medium storing a computer program, where the computer program, when executed by a computer, can implement any one of the methods. Brief Description of the Drawings

[0046] The drawings described herein are provided to further understand the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0047] Figure 1A A set of interface diagrams of the search interface for a mobile phone to enter the gallery application provided by an embodiment of the present application;

[0048] Figure 1B A schematic diagram of the search interface after clearing the search history provided by an embodiment of the present application;

[0049] Figure 1C A set of interface diagrams related to the search in the gallery application provided by an embodiment of the present application;

[0050] Figure 1D A schematic diagram of the search failure interface provided by an embodiment of the present application;

[0051] Figure 2A A schematic diagram of the structure of an electronic device provided by another embodiment of the present application;

[0052] Figure 2B A software structure block diagram of an electronic device provided by another embodiment of the present application;

[0053] Figure 3A Another set of interface diagrams related to the search in the gallery application provided by an embodiment of the present application;

[0054] Figure 3B A set of interface diagrams related to the search in the negative first screen provided by an embodiment of the present application;

[0055] Figure 3C A schematic diagram of the search result interface one provided by an embodiment of the present application;

[0056] Figure 3D A schematic diagram of the search result interface two provided by an embodiment of the present application;

[0057] Figure 3E A schematic diagram of the search result interface three provided by an embodiment of the present application;

[0058] Figure 3F A schematic diagram of the search result interface provided by an embodiment of the present application Figure Four ;

[0059] Figure 4 An interaction diagram of the visual media search method provided by an embodiment of the present application;

[0060] Figure 5Schematic flowchart of a visual media search method provided by an embodiment of the present application;

[0061] Figure 6 Schematic diagram of a search result interface provided by an embodiment of the present application Figure Five ;

[0062] Figure 7 Schematic flowchart of a filtering method based on semantic entities provided by an embodiment of the present application. Detailed implementation manners

[0063] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B; herein, "and / or" is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. These three situations. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0064] Hereinafter, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include one or more of such features.

[0065] First, the terms involved in the embodiments of the present application will be described. It can be understood that this description is for a clearer understanding of the embodiments of the present application and does not necessarily constitute a limitation to the embodiments of the present application.

[0066] Visual media: refers to pictures or videos.

[0067] Semantic entity: Named entity recognition technology can identify text and recognize entities with specific meanings in the text, such as person names (PER), place names (LOC), etc. In this solution, the entities with specific meanings identified are called semantic entities.

[0068] Visual content-related and visual content-unrelated: Visual content refers to the targets presented by visual media and their mutual relationships, etc. Computer vision can enable a computer to have capabilities similar to human vision, including perceiving, understanding, analyzing, and interpreting visual content. Currently, the Generative Pre-trained Transformer 4 (GPT-4) can support inputting an image into the model and then outputting human natural language describing the important information in the picture.

[0069] In the context of image search in this solution, the data that can only be obtained after the visual media file goes through the natural image understanding of the model is called "visual content-related". That is to say, visual content is the data that needs to be obtained through the natural image understanding model. In this solution, the data that is related to the visual media file and can be obtained without the image understanding ability of the model is called "visual content-unrelated", such as the shooting location, shooting time, name, file attributes, etc. that can be obtained and saved when the terminal device captures the visual media file.

[0070] For example, in the sentence "a photo taken in Beijing this year", "this year" (shooting time), "Beijing" (shooting location), and "photo" (file attribute) are all data that can be obtained and saved when the terminal device captures the visual media file. Therefore, "this year", "Beijing", and "photo" are unrelated to visual content; in the sentence "the sky taken in Beijing this year", "the sky" can only be obtained by the model's image understanding ability to understand the image or image frame of the visual media file. Therefore, "the sky" is visual content-related.

[0071] Text semantic vector: It can be obtained by sending the text into a text encoder. It is a vector that can represent the semantic features of the entire sentence. The text encoder can use models such as Transformer commonly used in natural language processing (NLP). This solution does not limit this here. In this solution, the text semantic vector obtained for a sentence is called a sentence semantic vector, the text semantic vector obtained for the semantic subject in a sentence is called a subject semantic vector, and the text semantic vector obtained for a label is called a label semantic vector.

[0072] Visual semantic vector: It can be obtained by sending the image or image frame of the visual media file into an image encoder. Commonly used CNN (Convolutional Neural Network) models or VIT (Vision Transformer) models can be used. This solution does not limit this here.

[0073] Density-based clustering algorithm: It describes the tightness of a sample set based on a set of neighborhoods. (The first parameter ∈, the second parameter MinPts) is used to describe the tightness of the sample distribution in the neighborhood. The first parameter ∈ is used to describe the neighborhood radius of a data point; the second parameter MinPts is used to describe the minimum number of data points in the neighborhood of a data point. Its representative algorithms include: DBSCAN (Density-Based Spatial Clustering of Application with Noise, density-based spatial clustering of applications with noise); the DBSCAN algorithm is a relatively representative density-based clustering algorithm that can divide regions with sufficient high density into clusters and can discover clusters of any shape in a spatial database with noise;

[0074] Vector similarity: It is used to describe the similarity between two vectors (for example: between a sentence semantic vector and a visual semantic vector). In the embodiments of the present application, the visual media that matches the search statement can be determined by comparing the similarity between the sentence semantic vector and the visual semantic vector. Generally, the vector similarity can be calculated through the cosine similarity calculation formula. Of course, it can also be calculated through other methods.

[0075] In the prior art, a mobile phone manages visual media files such as pictures and videos of users through a gallery application (hereinafter referred to as: visual media). Taking the example of a mobile phone taking a photo, after the mobile phone takes a photo, the gallery application can obtain and save attributes unrelated to the visual content such as the shooting location, shooting time, and photo name of the photo.

[0076] In practical applications, the photo can also be input into the natural picture understanding model in the state of the mobile phone being charged and the screen turned off, so that the natural picture understanding model generates and saves the label of the photo. This label can be regarded as an attribute related to the visual content of the photo. The label can be: "sky", "cat", "dog", etc. The gallery application can establish an index for the photo based on the attributes of the photo. After the index is established, the gallery application can provide corresponding search services to the user. Specifically, the user can search for pictures or videos by entering keywords in the gallery application. Exemplarily, the user can enter keywords such as "Beijing", "sky", "National Day" in the search box provided by the gallery application. The gallery application matches the keywords entered by the user with the indexes of visual media such as pictures and videos in the gallery application, and then obtains the search results.

[0077] The following describes the interface involved in the search process of the gallery application in the prior art with reference to the accompanying drawings:

[0078] As Figure 1AAs shown in (a) of Figure 1A , the mobile phone can display the main interface 101, which can also be called the desktop. The main interface 101 may include an icon 102 of the gallery application. The mobile phone can receive the operation of the user clicking on the icon 102. In response to this operation, the mobile phone can launch the gallery application and display the interface 103 as shown in

[0079] As shown in Figure 1A (b). Among them, the interface 103 may be an album interface. It should be noted that in response to the operation of the user clicking on the icon 102, the mobile phone can launch the gallery application and display the photo interface of the gallery. The photo interface includes thumbnails of the photos (i.e., pictures) in the gallery or a large picture of a certain photo. In the photo interface, in response to the operation of the user on the "Album" control, the above-mentioned album interface 103 is displayed.

[0080] As shown in Figure 1A (b), the interface 103 may include a search box 104. The mobile phone can receive the operation of the user clicking on the search box 104. In response to this operation, the mobile phone can display the interface 105 as shown in Figure 1A (c), and this interface 105 can be called the search interface. Among them, the interface 105 can display the classification information of the photos to the user. For example, in the interface 105, the mobile phone classifies the photos on the local machine according to time, people, and things, etc. For example, in the dimension of time, the mobile phone classifies the photos on the local machine according to the three time periods of "this month", "last month", and "this year" respectively. Among them, the "This Month" album includes the photos or videos taken by the mobile phone this month, the "Last Month" album includes the photos or videos taken by the mobile phone last month, and the "This Year" album includes the photos or videos taken by the mobile phone this year. In the dimension of people, the mobile phone classifies the photos on the local machine according to different people, such as the four different people in the interface 105. In the dimension of things, the mobile phone classifies and displays the photos on the local machine according to "scenery", "animals", "documents", and "buildings". It should be noted that the above classification dimensions can also be others, and no specific restrictions are made here. In the interface 105, the user can see this classification information without entering keywords.

[0081] Optionally, the interface 105 may further include a search history 107 and an option of "Clear" 108. The search history includes keywords that the user has entered, such as "flowers", "coffee", "cats", etc. The mobile phone can receive the operation of the user clicking on "Clear" 108, and in response to this operation, the mobile phone can clear the search history. After the mobile phone clears the search history, the keywords that the user has entered are no longer displayed on the search interface 105. For example, in response to the operation of the user clicking on "Clear" 108, as Figure 1B shown, the search history 107 and the option of "Clear" 108 are no longer displayed on the search interface 105, and the content displayed below moves up.

[0082] In response to the operation of the user entering the keyword "sky" on the interface 105, the mobile phone displays the interface 109 as shown in Figure 1C (a). As shown in Figure 1C (a), 100 photos related to "sky" and 32 photos related to photos containing the word "sky". Among them, the 100 photos related to "sky" can be recalled because the tags of these 100 photos match "sky"; the 32 photos related to photos containing the word "sky" can be recalled because through the OCR (Optical Character Recognition) technology, it is recognized that these 32 photos contain the word "sky". In practical applications, the mobile phone can also associate the keywords entered by the user to obtain associated words and perform a search based on the associated words.

[0083] The interface 109 also displays some search results of the keyword "sky" and a "More" option 110 corresponding to the search results of the keyword "sky". The mobile phone receives the click operation of the user on the "More" option 110 and displays the interface 111 as shown in Figure 1C (b). Among them, the interface 111 is used to display photos and videos in the search results of the keyword "sky". Optionally, the photos and videos can be classified and displayed according to time. In addition, the interface 111 also includes a return key 112 and a title 113. In response to the operation of the user on the return key 112, the mobile phone can redisplay the interface 109. The title 113 may include the keyword "sky".

[0084] That is to say, in the existing gallery applications, when a user enters simple keywords in the search box, such as: sky, Beijing, National Day, etc., corresponding search results can be obtained. However, since the number of photo attributes is simple and limited, and the mobile phone's ability to understand and associate search statements is also limited. When the user enters a relatively complex search statement in the search box, if the keywords in the search statement cannot match the attributes of the picture or the text in the picture, no photos can be found. That is to say, the existing gallery applications do not support the search function based on complex search statements. As Figure 1D shown, when the user enters a relatively complex search statement "warming oneself by the fire and brewing tea" in the search box of interface 114, the mobile phone cannot understand the associated words of "warming oneself by the fire and brewing tea". Since the photo has no attributes that can match "warming oneself by the fire and brewing tea" or its associated words, the search result shows "no pictures".

[0085] However, in practical applications, users have a strong demand for the function of searching for pictures based on complex search statements. This is because users can describe the pictures or videos they want more comprehensively through complex search statements, thereby achieving precise search. To meet this demand of users, an embodiment of this application provides a visual media search method. This method can be applied to an electronic device, and the electronic device can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and other terminal devices.

[0086] Exemplarily, Figure 2A shows a schematic structural diagram of an electronic device 200. The electronic device 100 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, antenna 1, antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone interface 270D, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.

[0087] Among them, the sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.

[0088] It can be understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on the electronic device 200. In other embodiments of this application, the electronic device 200 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0089] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0090] The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0091] A memory may also be provided in the processor 210 for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory may save the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0092] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0093] It can be understood that the interface connection relationships among the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0094] The electronic device 200 implements the display function through a GPU (Graphics Processing Unit), a display screen 294, and an application processor, etc. The GPU is connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change display information.

[0095] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 294, where N is a positive integer greater than 1.

[0096] The electronic device 200 can implement the shooting function through an ISP, a camera 293, a video codec, a GPU, a display screen 294, an application processor, etc.

[0097] The camera 293 is used to capture static images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In some embodiments, the electronic device 100 may include one or N cameras 293, where N is a positive integer greater than 1.

[0098] The video codec is used to compress or decompress digital videos. The electronic device 200 can support one or more video codecs. In this way, the electronic device 200 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0099] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously learn on its own. Through the NPU, applications such as intelligent cognition of the electronic device 200 can be realized, such as image recognition, face recognition, speech recognition, text understanding, etc.

[0100] The external memory interface 220 can be used to connect to an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 200. The external memory card communicates with the processor 210 through the external memory interface 220 to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.

[0101] The internal memory 221 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 221 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device 200 (such as audio data, phone book, etc.). In addition, the internal memory 221 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 210 executes various functional applications and data processing of the electronic device 200 by running the instructions stored in the internal memory 221 and / or the instructions stored in the memory provided in the processor.

[0102] Figure 2B It is a software structure block diagram of the electronic device 200 in the embodiment of the present application. The software system of the electronic device 200 can adopt a layered architecture. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. Taking the Android system as an example, in some embodiments, the Android system can be divided into four layers, from top to bottom are the application layer, the application framework layer, the Android runtime and the system library, and the kernel layer.

[0103] As Figure 2B shown, the application layer can include application programs such as a gallery service module, a search module, a multimodal understanding module, a natural language understanding module, and a camera application.

[0104] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.

[0105] The system library can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (e.g., OpenGL ES), 2D graphics engine (e.g., SGL), etc. Among them, the media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0106] The kernel layer is the layer between hardware and software.

[0107] Next, in combination with the scenario of capturing a photo, the working processes of the software and hardware of the electronic device 200 will be exemplarily described.

[0108] When the touch sensor 280K receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including information such as touch coordinates and the timestamp of the touch operation). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking the touch operation as a touch click operation and the control corresponding to the click operation being the control of the camera application icon as an example, the camera application calls the interface of the application framework layer to start the camera application, and then starts the camera driver by calling the kernel layer to capture a static image or video through the camera 293.

[0109] Taking the electronic device as a mobile phone as an example, the visual media search method provided by the embodiments of the present application will be introduced. The visual media search method provided by the embodiments of the present application can be applied in application programs such as the gallery application and the file management application.

[0110] Next, the interfaces and search logics involved in the visual media search method provided by the embodiments of the present application will be described in conjunction with the accompanying drawings.

[0111] As Figure 3A shown in (a) of [], the interface 301 displays search history 303 and a "Clear" option 304. The search history includes search statements that the user has entered, such as: "Watching the sunrise on the mountain top", "The Great Wall photographed in Beijing this year". Other content that can be displayed on the interface 301 can refer to the relevant display content of the above interface 305, which will not be elaborated here. Among them, the interface 301 can be called a search interface. The mobile phone can respond to the user's operation onFigure 1A Performing a click operation on the search box 104 in the interface 103 shown in (b) in displays the interface 301.

[0112] As Figure 3A shown in (b) in , in the interface 305 (i.e., the search result interface), the user enters the search statement "warming the tea around the stove" in the search box 306 of the interface 305, and the mobile phone searches for 239 pictures. The mobile phone displays some search results of the search statement "warming the tea around the stove" (e.g., thumbnails of 8 pictures) and the "More" option 307 corresponding to the search results of the search statement "warming the tea around the stove" on the interface 305. In response to the user's operation on the "More" option 307, the mobile phone displays the interface 308 shown in (c) in . The interface 308 can display the pictures in the search results of the search statement "warming the tea around the stove" in descending order according to the matching degree of the visual content of the pictures and the search statement "warming the tea around the stove" (specifically, it can be the similarity between the visual semantic vector of the picture and the sentence semantic vector of the search statement). The user can also perform an upward sliding operation on the interface 308 to view the pictures that have not been displayed. Figure 3A

[0113] Currently, the mobile phone also has a negative first screen, a pull-down search interface, etc. It can be understood that the negative first screen can be the leftmost split screen of the electronic device, which is used to provide functions such as search and quick services for the user. Among them, the negative first screen can also be used to display notification messages to be pushed to the user, such as application messages subscribed by the user, real-time hot search messages, segment selection, itinerary information, etc. The pull-down search interface is an interface displayed in response to the user's pull-down operation on the main interface. This interface is used to provide functions such as search and application suggestions for the user, and this interface and Figure 3B the interface 315 in can be the same interface.

[0114]

[0115] Figure 3B Figure 3B Exemplarily, referring to Figure 3B shown in (a) in , the mobile phone can receive a first operation performed by the user on the interface 309 (which can be called the desktop) of the mobile phone. Exemplarily, this first operation can be a rightward sliding operation as shown in (a) in . In response to this first operation, the mobile phone can display the negative first screen 310 shown in (b) in . Among them, the negative first screen 310 can include: a search box 311, quick services 312, default cards 313, recommended cards 314, etc. The quick services 312 can be quick access to a certain page or function of an application program, such as: scan code, payment code, ride code, etc.; the default cards can be: gallery cards, remaining battery cards, etc.; the recommended cards can be recommended application cards.​​​​

[0116] The mobile phone receives a click operation by the user on the search box 311 on the negative first screen 310, and displays the interface 315 shown in Figure 3B (c) below. The interface 315 may include: a search box 316, application suggestions, and a search history 217. The application suggestions include: icons of each application recommended for use. The interface 215 may also include: a search history 217 and its corresponding "Clear" option 318. In response to a trigger operation by the user on the "Clear" option 318, the search history 317 and the "Clear" option 318 are no longer displayed on the interface 315. In addition, a hot news title, such as: "Tianjin Marathon", may be displayed in the search box 316.

[0117] As Figure 3B (d) below shows the interface 319. The search box of the interface 319 displays the search statement "Great Wall photographed in Beijing during the National Day" entered by the user; a preview area 322 of the search results of the image gallery for the search statement "Great Wall photographed in Beijing during the National Day" and a corresponding "Search in the application" option 323 of the image gallery are also displayed in the interface 319. In response to a trigger operation by the user on the preview area 322, the mobile phone enters a photo details interface provided by the image gallery application for the user to flip through and view the search results of the search statement "Great Wall photographed in Beijing during the National Day". In response to a trigger operation by the user on the "Search in the application" option 323, the mobile phone displays the interface 324 provided by the image gallery application shown in Figure 3B (e) below. The interface 324 displays some search results of the search statement "Great Wall photographed in Beijing during the National Day" and a "More" option corresponding to the search results of the search statement "Great Wall photographed in Beijing during the National Day". In response to a trigger operation by the user on the "More" option, the mobile phone may display a search result details interface, and the search result details interface shows the pictures in the search results of the search statement "Great Wall photographed in Beijing during the National Day". An online search option 321 may also be displayed in the interface 319. In response to a trigger operation by the user on the online search option 321, the mobile phone displays a search web page and shows the online search results in the search web page.

[0118] As Figure 3CThe interface 325 shown. The user enters the search statement "The Great Wall photographed during last year's National Day" in the search box of the interface 325, and the mobile phone searches for 419 pictures. The mobile phone displays the thumbnails of each picture or video in the search results of the search statement "The Great Wall photographed during last year's National Day" on the interface 325. Among them, the shooting time of the picture or video corresponding to thumbnail A is 23:22 on September 30, 2022; the shooting time of the picture or video corresponding to thumbnail B is 00:19 on October 1, 2022. Among them, the time search range corresponding to National Day is from 00:00 on October 1, 2022 (i.e., the first time point) to 00:00 on October 8, 2022 (i.e., the second time point). The shooting time of the picture or video corresponding to thumbnail A is before 00:00 on October 1, 2022. The shooting time of the picture or video corresponding to thumbnail B is 00:19 on October 1, 2022, which is between 00:00 on October 1, 2022 and 00:00 on October 8, 2022. Among them, the shooting times of the pictures or videos corresponding to thumbnail A and thumbnail B are relatively close.

[0119] As Figure 3D For the interface 326 shown, for the search statement "The sky photographed during last year's National Day", the shooting time of the picture or video corresponding to the displayed thumbnail C is 22:19 on October 7, 2022; the shooting time of the picture or video corresponding to the displayed picture D is 01:24 on October 8, 2022. Among them, the time search range corresponding to National Day is from 00:00 on October 1, 2022 (i.e., the first time point) to 00:00 on October 8, 2022 (i.e., the second time point). The shooting time of the picture or video corresponding to thumbnail D is after 00:00 on October 8, 2022. The shooting time of the picture or video corresponding to thumbnail C is 00:19 on October 1, 2022, which is between 00:00 on October 1, 2022 and 00:00 on October 8, 2022. Among them, the shooting times of the pictures or videos corresponding to thumbnail C and thumbnail D are relatively close.

[0120] In practical applications, the search statement may include, in addition to the time search range (such as National Day), a first keyword (such as the Great Wall or the sky). Then, for the search statement, the visual media file corresponding to its search results matches the first keyword.

[0121] When the first keyword belongs to a semantic entity irrelevant to visual content (such as: location), the first keyword can be matched with the attributes of the visual media to determine the visual media file that matches the first keyword.

[0122] When the first keyword belongs to a semantic entity related to visual content (such as the Great Wall or the sky), match the first keyword with the tags of the visual media files to determine the visual media files whose visual content matches the first keyword; or, match the text semantic vector of the first keyword with the visual semantic vector of the visual media files to determine the visual media files whose visual content matches the first keyword (including the above-mentioned first, second, and third visual media files).

[0123] The above thumbnails are proportional thumbnails of the corresponding visual media files or proportional thumbnails of any image frame.

[0124] As Figure 3E Shown in the interface 327, in which, for the search statement "the sky photographed during last year's National Day", the two thumbnails E and F shown are sorted and displayed in descending order according to the matching degree between the visual content of their respective visual media files and the first keyword "the sky". Among them, the matching degree between the visual content of the visual media file corresponding to the thumbnail E and "the sky" is 0.87; the matching degree between the visual content of the visual media file corresponding to the thumbnail F and "the sky" is 0.76.

[0125] As Figure 3F Shown in the interface 328, in which, for the search statement "photos taken in Nankai District, Tianjin during last year's National Day", the two thumbnails J and H shown are sorted and displayed in descending order according to the matching degree between the location attributes of their respective visual media files and the first keyword "Nankai District, Tianjin". Among them, the matching degree between the location attribute of the visual media file corresponding to the thumbnail J and "Nankai District, Tianjin" is 0.87; the matching degree between the location attribute of the visual media file corresponding to the thumbnail H and "Nankai District, Tianjin" is 0.8.

[0126] Figure 4 This is the interaction diagram of the visual media search method provided by the embodiments of the present application. As Figure 4 Shown in the figure, the mobile phone is provided with: a gallery service module (i.e., a gallery application) 41, a search module 42, a multimodal understanding module 43, and a natural language understanding module 44.

[0127] As Figure 4 Shown in the figure, the visual media search method provided by the embodiments of the present application can be divided into two stages: an index construction stage and a search stage.

[0128] In the index construction stage, the following steps are included:

[0129] S401, Add and / or modify visual media and their attributes.

[0130] In the above S401, for the newly added visual media, the attributes that the gallery application can automatically generate and are unrelated to the visual content may include but are not limited to: the collection location, the collection time, and the name of the visual media. Taking a captured video or picture as an example, the collection location refers to the shooting location, and the collection time refers to the shooting time; taking a screenshot as an example, the collection location refers to the screenshot location, and the collection time refers to the screenshot time; taking a downloaded video or picture as an example, the collection location refers to the download location, and the collection time refers to the download time.

[0131] Users can add new visual media by means such as shooting, downloading, and taking screenshots. In addition, users can also modify the existing visual media. Such modifications include but are not limited to: beautification, custom naming, adding watermarks, etc.

[0132] S402. Store the visual media and its attributes.

[0133] The gallery service module 41 can, in response to the above-mentioned addition or modification operations, store the visual media and its attributes locally on the mobile phone. In actual applications, with the user's authorization, the mobile phone can store the locally stored visual media and its attributes in the cloud to relieve the storage pressure on the local mobile phone.

[0134] S403. Request visual semantic understanding of the visual media.

[0135] Since visual semantic understanding requires a large amount of computing resources, in order not to affect the user's use, the above step S403 can be executed when the mobile phone is in the state of being charged and the screen is off.

[0136] The gallery service module 41 can request the multimodal understanding module 43 to perform visual semantic understanding on the newly added or modified visual media to obtain the visual semantic vector of the visual media.

[0137] Among them, the multimodal understanding module 43 can perform visual semantic understanding on the visual media based on the multimodal model to obtain the visual semantic vector of the visual media.

[0138] Among them, the multimodal model is not only used for: performing visual semantic understanding on the visual media to obtain the visual semantic vector of the visual media; but also used for: performing semantic understanding on the search statement to obtain the sentence semantic vector of the search statement; performing semantic understanding on the rewritten search statement in the following text to obtain the sentence semantic vector of the rewritten search statement; performing semantic understanding on the semantic subject in the search statement to obtain the subject semantic vector of the semantic subject; performing semantic understanding on the label of the visual media to obtain the label semantic vector of the label. The multimodal model can be trained based on training samples.

[0139] Exemplarily, the multimodal model can specifically be CLIP (Contrastive Language-Image Pre-training). The CLIP model can map visual media and text (i.e., search statements, rewritten search statements, semantic entities, tags) into a unified vector space to understand the relationships between different modal resources visually and textually, and then be used for image retrieval. That is, in the embodiments of the present application, the visual media file and text can be specifically matched through the CLIP model.

[0140] The multimodal model can map visual media and text into vectors of the same dimension. That is to say, the dimension of the visual semantic vector of visual media is the same as that of the semantic vector of text (e.g., the sentence semantic vector of the search statement). The multimodal model includes: the above-mentioned image encoder and text encoder.

[0141] S404. Return the visual semantic vector of the visual media.

[0142] The multimodal understanding module 44 returns the visual semantic vector of the visual media to the gallery service module 41.

[0143] S405. Store the visual semantic vector of the visual media.

[0144] The gallery service module 41 can locally store the visual semantic vector of the visual media.

[0145] S406. Send the attribute information of the visual media and its visual semantic vector.

[0146] Exemplarily, the gallery service module 41 can store the visual semantic vector of the visual media returned by the multimodal understanding module 44, and then batch send the attributes of the visual media and its visual semantic vector to the search module 42 for the search module 42 to construct an index of the visual media.

[0147] S407. Construct an index

[0148] The index of the visual media constructed by the search module 42 can include: the attributes of the visual media and the visual semantic vector of the visual media.

[0149] In the search phase, the following steps are included:

[0150] S408. Input a search statement.

[0151] The user can input a search statement through the search interface provided by the gallery service module 41, for example: Figure 3A the interface 301 shown in (a) of Figure 3AIn the interface 305 shown in (b) thereof, enter "warming the tea around the stove" in the search box 306.

[0152] S409: Send the search statement.

[0153] After the gallery service module 41 receives the search statement input by the user, it sends the search statement to the search module 42 for searching.

[0154] S410: Request semantic subject recognition for the search statement.

[0155] The search module 42 requests the natural language understanding module 44 to perform semantic subject recognition on the search statement to obtain the semantic subjects included in the search statement. Among them, the natural language understanding module 44 performs semantic subject recognition based on the natural language understanding model. Specifically, named entity recognition technology (NER) can be used to perform semantic subject recognition on the search statement to obtain the semantic subjects included in the search statement. In the embodiments of the present application, the semantic subject can also be referred to as an entity.

[0156] Using named entity recognition technology, semantic subjects related to time, location, and tags in the search statement can be recognized. Among them, semantic subjects related to time and location are semantic subjects unrelated to visual content; the tags of visual media files are data that can only be obtained through the natural picture understanding of the model. Therefore, semantic subjects related to tags are semantic subjects related to visual content. In practical applications, based on practical experience, multiple tags that users are more concerned about can be counted, such as: "sky", "cat", "dog", "birthday", "child", etc. These tags are used to describe visual content. Subsequently, named entity recognition technology can match the keywords (or search terms) in the search statement with the multiple pre-set tags to determine whether the keyword belongs to the semantic subject related to the tag.

[0157] Exemplarily, using named entity recognition technology to perform semantic subject recognition on the search statement "the sky photographed in Beijing during the National Day", it is determined that "National Day" belongs to the semantic subject related to time, "Beijing" belongs to the semantic subject related to location, and "sky" belongs to the semantic subject related to tags.

[0158] S411: Return the semantic subject.

[0159] The natural language understanding module 44 returns the recognized semantic subject to the search module 42.

[0160] S412: Request semantic understanding of the search statement.

[0161] The search module 42 can send a search statement to the multimodal understanding module 43, and the multimodal understanding module 43 performs semantic understanding on the search statement to obtain the sentence semantic vector of the search statement (i.e., the first sentence semantic vector). For the specific semantic understanding process, reference can be made to the corresponding content in the above embodiments, which will not be elaborated here.

[0162] It should be additionally supplemented that when the search statement includes a semantic entity related to visual content (i.e., a semantic entity related to a label), the search module 42 can also send the semantic entity related to visual content to the multimodal understanding module 43, so that the multimodal understanding module 43 performs semantic understanding on the semantic entity related to visual content to obtain the entity semantic vector of this semantic entity. Continuing with the above example, "sky" belongs to the semantic entity related to visual content, and the multimodal understanding module 43 can perform semantic understanding on "sky" to obtain the entity semantic vector corresponding to "sky".

[0163] S413. Return the text vector.

[0164] The multimodal understanding module 43 can return the sentence semantic vector of the search statement to the search module 42.

[0165] S414. Perform recall respectively based on the attributes of the visual media and the visual semantic vector of the visual media.

[0166] The search module 42 includes different branches of search methods:

[0167] Exemplarily, the first branch is: performing recall based on the visual semantic vector of the visual media. Specifically, obtain the visual semantic vectors of each visual media among multiple visual media stored in the mobile phone; calculate the vector similarity between the sentence semantic vector of the search statement and the visual semantic vectors of each visual media; according to the vector similarity, determine M candidate visual media (i.e., M candidate visual media files) whose visual semantic vectors match the sentence semantic vector from the multiple visual media stored in the mobile phone; where M is an integer greater than or equal to 1; these M visual media can be used as the visual media set recalled by the first branch.

[0168] Exemplarily, use the visual media with a vector similarity greater than a preset similarity threshold as the visual media whose visual semantic vector matches the sentence semantic vector.

[0169] Exemplarily, sort the multiple visual media according to the vector similarity from high to low, and use the top F (F≥1) visual media as the visual media whose visual semantic vector matches the sentence semantic vector.

[0170] Exemplarily, Q (Q≥1) visual media with vector similarity greater than a preset similarity threshold are determined from multiple visual media stored via the mobile phone; if Q is greater than or equal to a preset quantity threshold D, these Q visual media can be sorted in descending order of vector similarity; the top D visual media in the sorting are used as the visual media whose visual semantic vectors match the semantic vector of this sentence; if Q is less than the preset quantity threshold D, these Q visual media can be directly used as the visual media whose visual semantic vectors match the semantic vector of this sentence.

[0171] It should be noted that the multiple visual media stored via the mobile phone may include: visual media stored locally by the mobile phone and / or visual media stored in the cloud by the mobile phone (for example: the cloud storage space applied for by the mobile phone). To protect user privacy, all the multiple visual media stored via the mobile phone are stored in the mobile phone.

[0172] It should be noted that in the recall process of the first branch, the attributes of the visual media are not understood, and only the overall visual semantic information of the visual media can be understood.

[0173] Exemplarily, the second branch is: recall based on the attributes of the visual media.

[0174] Specifically, obtain the attributes of each visual media among the multiple visual media stored via the mobile phone; match the semantic subject in the search statement with the attributes of each visual media to determine the visual media (i.e., the second visual media) that matches this semantic subject; use the visual media that matches this semantic subject as the visual media set recalled by the second branch. When the number of semantic subjects in the search statement is one, the visual media set recalled by the second branch includes the visual media that matches this one semantic subject; when the number of semantic subjects in the search statement is multiple, the visual media set recalled by the second branch includes the visual media that each of these multiple semantic subjects matches.

[0175] Exemplarily, for a semantic subject related to time, the corresponding time range (i.e., the time search range) of this semantic subject can be determined; match this time range with the acquisition time of each visual media as an attribute to determine the visual media whose acquisition time is within this time range; use the visual media whose acquisition time is within this time range as the visual media that matches this semantic subject. For example: the semantic subject is "National Day", its corresponding time range is "from October 1st to October 7th", the acquisition time of Picture 1 is "October 2nd", and the acquisition time of Picture 2 is "October 8th", then, according to the above matching method, Picture 1 matches the semantic subject "National Day", and Picture 2 does not match the semantic subject "National Day".

[0176] Exemplarily, for a semantic entity related to a location, the geographical range corresponding to the semantic entity (i.e., the geographical search range) can be determined; the geographical range is matched with the attribute of the collection location of each visual medium to determine the visual media whose collection location is within the geographical range; the visual media whose collection location is within the geographical range is used as the visual media matched by the semantic entity. For example: the semantic entity is "Beijing", its corresponding geographical range is the whole of Beijing, the collection location of Picture 3 is "Xicheng District, Beijing", and the collection location of Picture 4 is "Nankai District, Tianjin", then according to the above matching method, Picture 3 is matched with the semantic entity "Beijing", and Picture 4 is not matched with the semantic entity "Beijing".

[0177] Exemplarily, for a semantic entity related to a tag, the main semantic vector of the semantic entity and the tag semantic vectors of the tags of each visual medium can be obtained; the vector similarity between the main semantic vector of the semantic entity and the tag semantic vectors of the tags of each visual medium is calculated; according to the vector similarity, the visual media matched by the semantic entity is determined. For example: the semantic entity is "human cub", and Picture 5 has a tag of "child". Through calculation, it is found that the main semantic vector of "human cub" is similar to the tag semantic vector of "child", that is, Picture 5 is matched with the semantic entity "human cub". In practical applications, after obtaining the set of visual media recalled by the first branch and the set of visual media recalled by the second branch, the candidate set of visual media can be determined according to the set of visual media recalled by the first branch and the set of visual media recalled by the second branch. In an optional implementation manner, the union or intersection of the set of visual media recalled by the first branch and the set of visual media recalled by the second branch can be used as the candidate set of visual media.

[0178] In practical applications, when a user searches for pictures, etc. on a mobile phone, sometimes the user pays attention to the visual semantic information of the pictures, sometimes the user pays attention to the attribute information such as the shooting location and shooting time of the pictures, and sometimes the user pays attention to both. Exemplarily, when the user searches for "pictures taken today", the user pays attention to the shooting time of the pictures; when the user searches for "the sky taken today", the user not only pays attention to the shooting time of the pictures, but also pays attention to the visual semantics of the pictures, that is, whether the picture content is the sky; when the user searches for "pictures taken while walking in Beijing", the user pays attention to the shooting location of the pictures; when the user searches for "pictures of walking taken in Beijing", the user not only pays attention to the shooting location attribute of the pictures, but also pays attention to the visual semantics of the pictures, that is, whether the picture content is a picture of walking.

[0179] Taking the two search statements of "photos taken in Beijing this year" and "the sky taken in Beijing this year" as examples, referring to the foregoing introduction, the semantic proportion of visual content in the search statement of "the sky taken in Beijing this year" is greater than that in the search statement of "photos taken in Beijing this year". Obviously, for the search statement of "photos taken in Beijing this year", it is more appropriate to use the visual media set recalled by the second branch as the candidate visual media set. Subsequently, the intersection of the visual media matching "this year" and the visual media matching "Beijing" in the candidate visual media set can be used as the final search result. The collection locations of all visual media in the final search result are Beijing, and the collection time is this year, which meets the user's search requirements. If the visual media set recalled by the first branch is used as the candidate visual media set, the following situation may occur: There is a picture of "a girl holding a camera and taking pictures" in the visual media set recalled by the first branch, but the shooting location of this photo is Shanghai and the shooting time is last year. Since the word "taking pictures" exists in the search statement of "photos taken in Beijing this year", and the action of "taking pictures" exists in the picture of "a girl holding a camera and taking pictures", there is a certain similarity between the sentence semantic vector of the search statement of "photos taken in Beijing this year" and the visual semantic vector of the picture of "a girl holding a camera and taking pictures". That is to say, the picture of "a girl holding a camera and taking pictures" may be recalled. That is, when the semantic proportion of visual content is relatively low, it is not appropriate to use the visual media set recalled by the first branch as the candidate visual media set.

[0180] In order to solve the above problems, in an optional implementation manner, the semantic proportion related to visual content in the search statement can be determined; according to the semantic proportion related to visual content in the search statement, the first extraction ratio corresponding to the visual media set recalled by the first branch and the second extraction ratio corresponding to the visual media set recalled by the second branch are determined; when the semantic proportion is greater than or equal to the preset proportion threshold, the first extraction ratio is greater than the second extraction ratio. Exemplarily, if the semantic proportion is S, the first extraction ratio is: β*S, and the second extraction ratio is: 1-β*S, where the value of β can be set according to actual needs, and the embodiments of the present application do not make specific limitations on this.

[0181] In another alternative embodiment, the semantic proportion related to visual content in the search statement can be determined; when the semantic proportion related to visual content in the search statement is greater than or equal to a preset proportion threshold, the visual media set recalled by the first branch is used as the candidate visual media set; when the semantic proportion related to visual content in the search statement is less than the preset proportion threshold, the visual media set recalled by the second branch is used as the candidate visual media set. That the semantic proportion related to visual content in the search statement is greater than or equal to the preset proportion threshold indicates that the search statement belongs to a high-semantic search statement; that the semantic proportion related to visual content in the search statement is less than the preset proportion threshold indicates that the search statement belongs to a low-semantic search statement. The specific process will be described in detail in the embodiments below in conjunction with Figure 5 will be introduced in detail.

[0182] S415. Visual media filtering.

[0183] The search module 42 can perform semantic subject filtering, spatio-temporal filtering, and / or person relationship filtering on the candidate visual media set. The specific filtering method will be described in detail in the embodiments below in conjunction with Figure 5 will be introduced in detail.

[0184] S416. Visual media ranking.

[0185] The search module 42 ranks the candidate visual media set.

[0186] The specific ranking method will also be described in detail in the following embodiments.

[0187] S417. Return search results.

[0188] The search module 42 sends the search results obtained after ranking to the gallery service module 41.

[0189] Specifically, the search module 42 can select the top P (P≥1) visual media and their ranking information as the final search results.

[0190] S418. Display search results.

[0191] The gallery service module 41 can display the search results to the user. For example: through Figure 3A interface 305 in Figure 3A interface 308 in

[0192] such as Figure 3A interface 305 in Figure 3AThe interface 308 therein. In one example, the display order of pictures in the search results is related to the matching degree between the pictures and the search results. For example, the matching degree between the pictures displayed earlier and the search statement is greater than or equal to the matching degree between the pictures displayed later and the search statement. The calculation of the matching degree and the sorting method will be introduced in detail in the following embodiments.

[0193] The following will combine Figure 5 to introduce the search process executed by the search module 42 of this application in detail:

[0194] 501. Receive the search statement.

[0195] 502. Execute the search process corresponding to the search statement based on the visual semantic vectors of the visual media.

[0196] The visual media set recalled by the first branch obtained by executing step 502 can be specifically referred to the search process corresponding to the first branch above.

[0197] Exemplarily, assume that the multimodal model can encode the visual media and the search statement into a k-dimensional vector space. The search statement is Q, and the corresponding sentence semantic vector V Q ={α z}, z = 1, 2,..., Z, and a total of M picture sets R are recalled, and their vector representations are:

[0198]

[0199] Among them, the i-th row corresponds to the visual semantic vector of the i-th picture, and the value range of i is [1, Mm].

[0200] 503. Semantic entity recognition.

[0201] For the specific process of semantic entity recognition of the search statement, reference can be made to the corresponding content in the above embodiments, and details will not be repeated here.

[0202] 504. Perform a search based on the attributes of the visual media.

[0203] Specifically, match the semantic entities irrelevant to the time content in the search statement with the attributes of the visual media to obtain the visual media set recalled by the second branch. For details, reference can be made to the search process corresponding to the second branch above.

[0204] 505. Rewrite the search statement.

[0205] Specifically, delete the semantic entities irrelevant to the visual content in the search statement to obtain the rewritten search statement. The semantic entities irrelevant to the visual content specifically refer to: semantic entities related to time and semantic entities related to location.

[0206] Exemplary: The search statement is "the sky photographed this year", where "today" is a semantic subject related to time, and the rewritten search statement is "the photographed sky".

[0207] In practical applications, after deleting the semantic subjects unrelated to visual content, there may be some redundant stop words. For example, for the search statement "the sky photographed in Beijing this year", where "this year" is a semantic subject related to time and "Beijing" is a semantic subject related to location, after deleting "this year" and "Beijing", the stop word "in" becomes a redundant word and thus also needs to be deleted. Specifically, delete the semantic subjects unrelated to visual content in the search statement and their related stop words to obtain the rewritten search statement. Exemplarily, the rewritten search statement corresponding to the search statement "the sky photographed in Beijing this year" is "the photographed sky".

[0208] It should be noted that there is no order restriction in the execution of the above steps 502, 503, and 505. In an optional example, to improve efficiency, these three steps can be executed simultaneously.

[0209] 506. Perform the search process corresponding to the rewritten search statement based on the visual semantic vector of the visual media.

[0210] Performing step 506 obtains the set of visual media recalled by the third branch.

[0211] Specifically, obtain the visual semantic vectors of each visual media among the multiple visual media stored in the mobile phone; calculate the vector similarity between the sentence semantic vector of the rewritten search statement (i.e., the second sentence semantic vector) and the visual semantic vectors of each visual media; according to the vector similarity, determine multiple visual media (i.e., reference visual media) whose visual semantic vectors match the sentence semantic vector from the multiple visual media stored in the mobile phone; these multiple visual media can be used as the set of visual media recalled by the third branch.

[0212] Among them, the specific implementation process of the step of "determining multiple visual media whose visual semantic vectors match the sentence semantic vector from the multiple visual media stored in the mobile phone according to the vector similarity" can refer to the corresponding content in the above embodiments and will not be elaborated here.

[0213] Exemplarily, the rewritten search statement is Q′, and the corresponding vector V Q′ ={α′ z}, z = 1, 2,..., Z, and a total of M′ picture sets R′ are recalled, and their vector representation is:

[0214]

[0215] Among them, the i-th row corresponds to the visual semantic vector of the i-th picture, and the value range of i is [1, M′].

[0216] 507. Calculate the semantic proportion related to visual content in the search statement.

[0217] The following will introduce a method for determining the semantic proportion related to visual content:

[0218] 5071. Determine the representative visual semantic vector V I (i.e., the first representative visual semantic vector) corresponding to the visual media set recalled by the first branch and the representative visual semantic vector V I′ (i.e., the second representative visual semantic vector) corresponding to the visual media set recalled by the third branch.

[0219] 5072. Determine the sentence semantic vector V Q and the difference from the representative visual semantic vector V I to obtain the difference vector (V Q -V I )(i.e., the first difference vector).

[0220] 5073. Determine the sentence semantic vector V Q′ and the difference from the representative visual semantic vector V I′ to obtain the difference vector (V Q′ -V I′ )(i.e., the second difference vector).

[0221] 5074. Determine the semantic proportion related to visual content in the search statement according to the vector similarity between the difference vector (V Q -V I ) and the difference vector (V Q′ -V I′ ).

[0222] Among them, the semantic proportion related to visual content is positively correlated with this vector similarity.

[0223] In the above 5071, in one example, the vector similarity between the visual semantic vector of each visual media in the visual media set recalled by the first branch and the sentence semantic vector of the search statement can be obtained; the visual media in the visual media set recalled by the first branch are sorted according to the vector similarity from high to low; the average vector of the visual semantic vectors of the top T (T≥1) visual media with higher rankings is used as the representative visual semantic vector V I .

[0224] Continuing with the above example, the representative visual semantic vector V I is:

[0225]

[0226] The vector similarity between each visual semantic vector of the visual media recalled by the third branch and the sentence semantic vector of the rewritten search statement can be obtained; the visual media in the visual media set recalled by the third branch are sorted according to the vector similarity from high to low; the average vector of the visual semantic vectors of the top H (H≥1) visual media in the sorting is used as the representative visual semantic vector V I′ 。

[0227] Continuing with the above example: The representative visual semantic vector V I′ is:

[0228]

[0229] The values of H and T above can be the same or different, and the embodiments of the present application do not make specific limitations on this. In another example, according to the vector similarity between the visual semantic vector of each visual media in the visual media set recalled by the first branch and the sentence semantic vector of the search statement, the visual media in the visual media set recalled by the first branch can be clustered using a clustering algorithm to obtain the clustering center point; the visual semantic vector of the clustering center point is used as the first representative visual semantic vector.

[0230] According to the vector similarity between the visual semantic vector of each visual media in the visual media set recalled by the third branch and the second sentence semantic vector of the rewritten search statement, the visual media in the visual media set recalled by the third branch can be clustered using a clustering algorithm to obtain the clustering center point; the visual semantic vector of the clustering center point is used as the second representative visual semantic vector.

[0231] In the embodiments of the present application, the representative visual semantic vector is used to represent the visual semantics of the entire set. The specific clustering algorithm can be selected according to actual needs, and the embodiments of the present application do not make any limitations on this.

[0232] In the above 5074, in an optional embodiment, the vector similarity between the difference vector (V Q -V I ) and the difference vector (V Q′ -V I′ ) can be directly used as the semantic proportion related to the visual content in the search statement.

[0233] Among them, the vector similarity between the difference vector (V Q -V I ) and the difference vector (V Q′ -V I′ ), the larger it is, the greater the semantic proportion of the visual content in the search statement; on the contrary, it means that the semantic proportion of the visual content in the search statement is smaller.

[0234] It should be noted that in the embodiments of the present application, the calculation method of the semantic proportion related to visual content in the search statement draws on the word analogy feature of the distribution representation vector, that is, the additivity of word meanings is directly reflected in the additivity of the distribution representation vector.

[0235] When the semantic proportion related to visual content in the search statement is less than the preset proportion threshold, step 508 is executed subsequently.

[0236] When the semantic proportion related to visual content in the search statement is greater than or equal to the preset proportion threshold, steps 509 and 510 are executed subsequently.

[0237] 508. Determine the visual media recalled by the second branch as the candidate set.

[0238] 509. Determine the visual media recalled by the first branch as the candidate set.

[0239] 510. Visual media filtering.

[0240] To improve the search accuracy, one or more of the following processing operations can also be performed on the visual media recalled by the first branch: semantic entity filtering, spatio-temporal filtering, person name filtering, and person relationship filtering. Among them, semantic entity filtering refers to filtering the visual media recalled by the first branch based on the matching degree between the visual media recalled by the first branch and the semantic entity related to visual content in the search statement. Spatio-temporal filtering includes: time filtering and space (i.e., location) filtering. Person name filtering refers to filtering the visual media recalled by the first branch based on the matching degree between the person name attribute of the visual media recalled by the first branch and the semantic entity related to the person name in the search statement. Person relationship filtering refers to filtering the visual media recalled by the first branch based on the matching degree between the person relationship attribute of the visual media recalled by the first branch and the semantic entity related to the person relationship in the search statement. In practical applications, in addition to the above-mentioned attributes such as the collection location, collection time, and visual media name stored in the mobile phone, the visual media may also include the person name attribute and person relationship attribute manually input by the user. Therefore, in practical applications, the named entity recognition technology can also be used to match the keywords in the search statement with a variety of preset person relationships to determine whether the keyword belongs to the semantic entity related to the person relationship.

[0241] When multiple processing operations such as semantic entity filtering, spatio-temporal filtering, person name filtering, and person relationship filtering need to be performed on the candidate set recalled by the first branch, the execution order of these multiple processing operations can be set according to actual needs, and the embodiments of the present application do not make specific limitations on this.

[0242] Exemplarily, as Figure 5 shown, the visual media filtering includes:

[0243] 5101. Semantic entity filtering.

[0244] 5102. Spatiotemporal filtering.

[0245] 5103. Character relationship filtering.

[0246] 5104. Person name filtering.

[0247] In the above 5101, generally, when a user searches, the images they hope to recall contain some visual content described in their search statement, such as: sky, puppy, child, etc. That is to say, the semantic entities related to visual content in the search statement represent the visual focus that the user pays attention to during the search.

[0248] However, when the first branch conducts the recall, it considers the matching degree between the visual content of visual media such as images and videos and the entire search statement of the user, without fully considering the role of certain specific semantic entities (i.e., the semantic entities related to visual content) in the search statement in the mobile phone gallery search scenario, resulting in the recall of some inaccurate photos.

[0249] For example: When the user searches for "photos of the Great Wall taken on National Day", the semantic entity related to visual content among them includes "the Great Wall". When the first branch makes the match, it considers the matching degree between the entire search statement and the visual content of the image, which may lead to the first branch recalling images that contain the "taking" behavior but do not contain the "Great Wall". This is because this image contains the "taking" behavior and the search statement contains the word "taking". That is to say, there is a certain similarity between the visual semantic vector of the image and the sentence semantic vector of the search statement. Then, this image has the possibility of being recalled.

[0250] Another example: When the user searches for "last year when the child had a birthday holding a cake", the semantic entities related to visual content among them include "child", "birthday", and "cake". When the first branch makes the match, it considers the matching degree between the entire search statement and the image, which may lead to the model recalling images of an adult having a birthday holding a cake. This is because the similarity between the visual semantic vector of the image of an adult having a birthday holding a cake and the sentence semantic vector of "last year when the child had a birthday holding a cake" is very high.

[0251] Therefore, in order to further improve the accuracy of searching for visual media on mobile phones, on the basis of the first branch recalling visual media, the "semantic entity" can be strengthened, that is, the results recalled by the first branch can be fine-tuned by means of the matching degree between the visual media and the "semantic entity".

[0252] Specifically, as Figure 7As shown, the "semantic entity filtering" in 5101 above may include the following steps:

[0253] 5101a. Determine the dimension to be matched according to the semantic entity related to the visual content in the search statement.

[0254] 5101b. Filter the M candidate visual media according to the matching degree of the visual content of the M candidate visual media in the dimension to be matched.

[0255] In 5101a above, in one example, the semantic entity related to the visual content is used as the dimension to be matched. When there are multiple semantic entities related to the visual content in the search statement, these multiple semantic entities are respectively used as different dimensions to be matched to obtain multiple dimensions to be matched. Among them, the number of multiple dimensions to be matched is the same as the number of semantic entities related to the visual content.

[0256] Exemplarily, for the search statement "Last year, the child had a birthday holding a cake", the semantic entities related to the visual content are "child", "birthday", and "cake". Then, the corresponding three dimensions to be matched are: "child", "birthday", and "cake".

[0257] In practical applications, in addition to using the semantic entity as the dimension to be matched, the search statement itself can also be used as the dimension to be matched. In this way, when performing semantic entity filtering, the matching situation between the visual content of the candidate visual media and the entire search statement can also be considered to improve the rationality of filtering. Specifically, the semantic entity related to the visual content and the search statement can be respectively used as different dimensions to be matched to obtain multiple dimensions to be matched. When there are multiple semantic entities related to the visual content in the search statement, these multiple semantic entities and the search statement are respectively used as different dimensions to be matched. Among them, the number of multiple dimensions to be matched is one more than the number of semantic entities related to the visual content.

[0258] Exemplarily, for the search statement "Last year, the child had a birthday holding a cake", the semantic entities related to the visual content are "child", "birthday", and "cake". Then, the corresponding four dimensions to be matched are: "child", "birthday", "cake", and "Last year, the child had a birthday holding a cake".

[0259] In the above 5101b, the matching degree of the visual content of the candidate visual media in the dimension to be matched refers to the matching degree between the visual content of the candidate visual media and the dimension to be matched. The vector similarity between the visual semantic vector of the candidate visual media and the semantic vector to be matched in the dimension to be matched can be determined; among them, when the dimension to be matched is the semantic subject related to the visual content, the semantic vector to be matched in this dimension to be matched is the subject semantic vector of this semantic subject; when the dimension to be matched is a search statement, the semantic vector to be matched in this dimension to be matched is the sentence semantic vector of this search statement; according to this vector similarity, the matching degree of the visual content of the candidate visual media in the dimension to be matched is determined. Among them, the matching degree is positively correlated with the vector similarity.

[0260] When the number of dimensions to be matched is one, the candidate visual media with a matching degree less than or equal to the preset matching degree threshold can be filtered out according to the matching degrees of the visual contents of multiple candidate visual media in the dimension to be matched.

[0261] For example: the search statement is "photos of the sky taken on National Day", and the only semantic subject related to the visual content is "sky"; then, the recalled pictures that contain the "taking" behavior but do not contain "sky" have a relatively low matching degree with "sky" and will be filtered out.

[0262] When the number of dimensions to be matched is multiple, for each candidate visual media, the comprehensive matching degree of this candidate visual media is determined according to the matching degrees of the visual content of this candidate visual media in multiple dimensions to be matched; multiple candidate visual media are filtered according to the comprehensive matching degree of each candidate visual media. Specifically, from M candidate visual media, multiple first visual media with a comprehensive matching degree greater than or equal to the preset matching degree threshold (that is, meeting the preset requirements) are determined, which is equivalent to filtering out the candidate visual media with a comprehensive matching degree less than the preset matching degree threshold; or, the M candidate visual media are sorted in descending order of the comprehensive matching degree, and the top Z′ (Z′≥1) candidate visual media (that is, meeting the preset requirements) are used as multiple first visual media, which is equivalent to filtering out the (M-Z′) candidate visual media ranked behind.

[0263] In an optional implementation manner, any one of the following three methods can be used to determine the comprehensive matching degree of the candidate visual media:

[0264] Method 1: Sum the matching degrees of the visual content of the candidate visual media in multiple dimensions to be matched to obtain the comprehensive matching degree of the candidate visual media.

[0265] Method 2: Perform a weighted sum of the matching degrees of the visual content of the candidate visual media in multiple dimensions to be matched to obtain the comprehensive matching degree of the candidate visual media.

[0266] Among them, the weights of multiple dimensions to be matched can be configured by the user in advance.

[0267] Method 3: Use a machine learning model to determine the comprehensive matching degree of candidate visual media according to the matching degrees of the visual content of the candidate visual media on multiple dimensions to be matched.

[0268] Among them, the machine learning model needs to be trained based on a data set, and the purpose of its training is essentially to learn the weights of each dimension to be matched.

[0269] In the above Method 1, the contribution degrees of the matching degrees in different dimensions to the comprehensive matching degree are not distinguished, which may lead to the numerical values of the comprehensive matching degrees of multiple candidate visual media calculated finally being relatively close or equal, and then it is impossible to screen multiple candidate visual media.

[0270] Exemplarily, assume that there are Picture 1, Picture 2, and Picture 3 among multiple candidate visual media, and there are Dimension A, Dimension B, and Dimension C among multiple dimensions to be matched. Calculate the matching degrees of the visual content of each picture on each dimension respectively, and the results are shown in Table 1.

[0271] Table 1:

[0272] Dimension A Dimension B Dimension C Picture 1 0.30 0.35 0.55 Picture 2 0.40 0.38 0.42 Picture 3 0.38 0.37 0.45

[0273] If calculated according to Method 1, the comprehensive matching scores of Picture 1, Picture 2, and Picture 3 are all 1.2, which will lead to the inability to screen the three pictures.

[0274] In the above Method 2, when there are too many semantic entities related to visual content to be concerned about (that is, the number of preset multiple tags is too large), it is difficult to accurately configure the weights of different dimensions.

[0275] In the above Method 3, the construction of the data set is inseparable from user data. However, user data belongs to user privacy content, and users do not want their data to be reported to the cloud side.

[0276] In an optional implementation manner, to solve the above problems, the following steps can be adopted to determine the comprehensive matching degree:

[0277] S51: Determine the weights of N dimensions to be matched respectively.

[0278] Among them, N is an integer greater than 1.

[0279] Among them, the weight of the j-th dimension to be matched is positively correlated with the degree of variation of the matching degree of the visual content of the M candidate visual media files on the j-th dimension; j is an integer, and the value of j ranges from 1 to N in sequence;

[0280] S52. According to the weights of the N dimensions to be matched, perform a weighted sum of the degrees of match of the visual content of the i-th candidate visual media file on each dimension to be matched among the N dimensions to be matched, so as to obtain the comprehensive degree of match of the i-th candidate visual media file.

[0281] Wherein, i is an integer, and the values of i sequentially range from 1 to M.

[0282] Taking the search statement "Last year, the child held a cake on his birthday" as an example, the semantic entities related to the visual content include: child, birthday, and cake. If multiple recalled photos all contain a cake, then the degree of variation of the degrees of match of the visual content of the multiple recalled photos on the dimension of "cake" will be relatively small, and the weight corresponding to the dimension of "cake" will be relatively small; if some of the multiple recalled photos contain a child and some do not, then the degree of variation of the degrees of match of the visual content of the multiple recalled photos on the dimension of "child" will be relatively large, and the weight corresponding to the dimension of "child" will be relatively large.

[0283] In one example, for each dimension to be matched, according to the information entropy of the degrees of match of the visual content of multiple candidate visual media on this dimension to be matched, determine the degree of variation of the degrees of match of the visual content of multiple candidate visual media on this dimension to be matched; wherein, the degree of variation is inversely proportional to the information entropy. The calculation method of information entropy will be introduced in detail in the following embodiments.

[0284] In this embodiment, the information entropy is used to measure the degree of variation of the degrees of match under each dimension to be matched, and based on this, the weight corresponding to this dimension to be matched is determined.

[0285] In order to ensure that the degrees of match of the visual content of candidate visual media on different dimensions to be matched have a unified dimension, a normalization processing step can be performed. Specifically, for each candidate visual media, according to the vector similarity between the visual semantic vector of the candidate visual media and the semantic vector to be matched of the dimension to be matched, determine the initial degree of match of the candidate visual media on the dimension to be matched. Exemplarily, the vector similarity between the visual semantic vector of the candidate visual media and the semantic vector to be matched of the dimension to be matched can be used as the initial degree of match of the candidate visual media on the dimension to be matched. Perform normalization processing on the initial degrees of match of the visual content of multiple candidate visual media on the dimension to be matched to obtain the degrees of match of the visual content of multiple candidate visual media on this dimension to be matched.

[0286] The normalization process and the calculation processes of weights and comprehensive degrees of match will be introduced in detail below:

[0287] Suppose there are m candidate visual media and n dimensions to be matched. The initial matching degrees of the visual content of the m candidate visual media on each dimension to be matched among the n dimensions to be matched can be regarded as a data matrix:

[0288] X = (x ij ) m*n , i = 1, 2, …, m; j = 1, 2, …, n (5)

[0289] where x ij is the initial matching degree of the visual content of the i-th candidate visual media on the j-th dimension to be matched.

[0290] Continuing with the above example, there are multiple candidate visual media including Picture 1, Picture 2, and Picture 3, and multiple dimensions to be matched including Dimension A, Dimension B, and Dimension C. That is, m is 3 and n is 3.

[0291] Step 1: Perform normalization processing on the above data matrix.

[0292] The normalized matrix is:

[0293] R = (r ij ) m*n , i = 1, 2, …, m; j = 1, 2, …, n (6)

[0294] where

[0295]

[0296] where max(x j ) refers to the maximum value among the initial matching degrees of the visual content of multiple candidate visual media on the j-th dimension to be matched; min(x j ) refers to the minimum value among the initial matching degrees of the visual content of multiple candidate visual media on the j-th dimension to be matched.

[0297] In practical applications, other normalization methods can also be used, and the embodiments of this application do not make specific limitations in this regard.

[0298] Exemplarily, the results obtained by normalizing the example data in Table 1 above are shown in Table 2.

[0299] Table 2:

[0300] Dimension A Dimension B Dimension C Picture 1 0.00 0.00 1.00 Picture 2 1.00 1.00 0.00 Picture 3 0.80 0.67 0.23

[0301] Step 2: Calculate the information entropy corresponding to each dimension to be matched.

[0302] Formulas (8) and (9) can be used for calculation:

[0303]

[0304] Among them,

[0305]

[0306] Among them, e j refers to the information entropy corresponding to the j-th dimension to be matched.

[0307] Among them, since the domain of the ln(x) function is x > 0. In actual calculations, to avoid the situation where p ij in ln(p ij ) takes 0, ln(p ij ) in formula (8) can be replaced by ln(p ij +α), where α << 0.001.

[0308] The smaller the information entropy corresponding to the dimension to be matched, the greater the degree of variation in the matching degree of the visual content of multiple candidate visual media in this dimension to be matched, and the greater the amount of information provided. It can be considered that the role played by this dimension to be matched in the comprehensive evaluation is also greater.

[0309] Exemplarily, for the example data in Table 2 above, the information entropy calculated according to the above formula (8) and formula (9) is shown in Table 3:

[0310] Table 3:

[0311] Dimension A Dimension B Dimension C Information Entropy 0.69 0.67 0.48

[0312] Step 3: Calculate the weights corresponding to each dimension to be matched.

[0313] The formula (10) can be used to calculate the weights corresponding to each dimension to be matched:

[0314]

[0315] Among them, d j refers to the weight corresponding to the j-th dimension to be matched.

[0316] It can be seen that the above formula (10) is a monotonically decreasing function of the information entropy. In this embodiment, the smaller the information entropy, the greater the degree of variation; the greater the degree of variation, the greater the weight. That is to say, the weight is negatively correlated with the information entropy.

[0317] To ensure that the sum of the weights corresponding to multiple dimensions to be matched is 1, the formula (11) can be used for the following calculation to obtain the final weight w j :

[0318]

[0319] Among them, w jRefers to the final weight corresponding to the j-th dimension to be matched.

[0320] In an alternative embodiment, other monotonically decreasing functions may also be used to calculate the above weights, and the embodiments of the present application do not make specific limitations thereto.

[0321] Exemplarily, for the example data in Table 3 above, the weights of different dimensions to be matched can be calculated according to the above formulas (10) and (11), as shown in Table 4:

[0322] Table 4:

[0323] Dimension A Dimension B Dimension C Weight 0.28 0.29 0.42

[0324] It should be noted that since rounding is introduced in the process of calculating the information entropy, the sum of the three dimensions in Table 4 above is not 1.

[0325] Step 4: Weighted summation.

[0326] Perform a weighted summation on the normalized matching degrees of any candidate visual media to obtain the comprehensive matching degree of the candidate visual media, which can be specifically calculated using formula (12):

[0327]

[0328] where s i Refers to the comprehensive matching degree of the i-th visual media.

[0329] Exemplarily, for the example data in Tables 2 and 4 above, the comprehensive matching degrees of different pictures are calculated using the above formula (12), as shown in Table 5:

[0330] Table 5:

[0331] Picture 1 Picture 2 Picture 3 Comprehensive Matching Degree 0.42 0.58 0.52

[0332] In the above 5102, when the first branch recalls visual media, it only considers the visual semantic information of the visual media, without considering the attribute information such as the time and location of the visual media. Therefore, it is necessary to perform time filtering, location filtering, etc. on the visual media recalled by the first branch. Specifically, when the user performs semantic search, if the search statement contains time information, pictures that do not meet the time limit need to be filtered out.

[0333] However, the user's description form of time information is rich and diverse and is fuzzy. For example, when the user searches for "the sky photographed in the afternoon", the "afternoon" in the user's search statement has no clear and standardized definition. When the user uses a fuzzy time expression in the search statement, the time window used for time filtering will affect the user experience.

[0334] Exemplary: The user starts traveling to other places on September 30, 2022 and arrives home on October 9, 2022. During this period, the user takes a lot of photos. One day in 2023, the user wants to view the beautiful scenery captured during this trip. Then the user is very likely to enter the search statement "Scenery captured during last year's National Day holiday". If directly based on "last year's National Day" in the search statement, the time window used for time filtering is set to "from October 1, 2022 to October 7, 2022", then the scenic photos captured by the user on September 30, 2022, October 8, 2022, and October 9, 2022 will be filtered out, which obviously does not meet the user's expectations. To improve the rationality of time filtering, the embodiments of the present application provide a new time filtering method. Specifically, using a clustering algorithm, multiple candidate visual media are clustered according to the acquisition time of each of the multiple candidate visual media, and K (K≥1) clustering clusters are obtained.

[0335] In this way, pictures of the same series with relatively close acquisition times can be grouped into the same clustering cluster.

[0336] Continuing with the above example, the natural scenery picture captured by the user on September 30, 2022 and the natural scenery picture captured by the user on October 1, 2022 have relatively close shooting times and are grouped into the same clustering cluster through the above clustering algorithm.

[0337] The above clustering algorithm may include but is not limited to: K-Means clustering algorithm, Mean shift clustering algorithm, and density-based clustering algorithm.

[0338] Taking the density-based clustering algorithm as an example, in the scenario of semantic search containing time information, it is unreasonable to set fixed first parameter ∈ and second parameter MinPts. For example, when the user searches for "Photos of going out to play in 2022", the time range is 1 whole year; while when the user searches for "Photos of going out to play in the morning", the time range is several hours. The same first parameter ∈ and second parameter MinPts should not be set in these two cases. To improve the rationality of clustering, the following steps can be used to determine the first parameter ∈ and the second parameter MinPts:

[0339] 51021. Determine the first parameter involved in the density-based clustering algorithm according to the time search range included in the search statement.

[0340] Among them, the first parameter is positively correlated with the duration corresponding to the time search range.

[0341] The time search range can be determined according to the time-related semantic entity in the search statement. Specifically, the time range corresponding to the time-related semantic entity (i.e., the time search range) can be returned through a mapping table. Among them, the mapping table can be constructed in advance as needed, and the specific form is not specifically limited in the embodiments of the present application.

[0342] Exemplarily, the time-related semantic entity extracted from the search statement "photos taken in spring" is "spring". Querying the mapping table, the corresponding time range is obtained as: February 1 - May 30; the time-related semantic entity extracted from the search statement "the sky taken in the morning" is "morning". Querying the mapping table, the corresponding event range is obtained as: 7:00 - 12:00; the time-related semantic entity extracted from the search statement "photos of going out for fun during the National Day" is "National Day". Querying the mapping table, the corresponding time range is obtained as: October 1 - October 7; the time-related entity extracted from the search statement "the sky taken in Beijing this year" is "this year". Querying the mapping table, the corresponding time range is obtained as "January 1, 2023 to December 31, 2023".

[0343] The first parameter can be determined according to the duration corresponding to the time search range. Exemplarily, the start time of the time search range (which can be understood as the start timestamp) is T start and the end time (which can be understood as the end timestamp) is T end , and the duration of the time search range is: T end -T start , and the following formula can be used to calculate the first parameter:

[0344] ∈=α*(T end -T start ) (13)

[0345] Among them, α is a coefficient that can be adjusted manually, and its size can be set according to actual needs, which is not specifically limited in the present application.

[0346] 51022. Determine the second parameter involved in the density-based clustering algorithm according to the ratio of the number of multiple candidate visual media to the time search range.

[0347] The following formula can be used to calculate the second parameter:

[0348]

[0349] Among them, N is the total number of visual media whose acquisition time is within the time search range in the candidate set, N≥1; among them, n is the number of times the time search range repeats between T 1 and T 2 . T 1is the acquisition time of the earliest acquired visual media among multiple visual media stored via the mobile phone; T 2 is the acquisition time of the latest acquired visual media among multiple visual media stored via the mobile phone. Exemplarily, the search statement is "the sky photographed during the National Day", and its time search range is from October 1st to October 7th, with a duration of 7 days; the acquisition time of the earliest acquired visual media stored in the mobile phone is August 1st, 2020; the acquisition time of the latest acquired visual media stored in the mobile phone is October 20th, 2023; then, from August 1st, 2020 to October 20th, 2023, the time search range from October 1st to October 7th repeats 4 times (that is, once a year).

[0350] where, (T end - T start ) * n can be understood as the total duration corresponding to the time search range.

[0351] where, β is a coefficient that can be adjusted manually, and its magnitude can be set according to actual needs. This application does not make specific limitations on this.

[0352] In this embodiment, according to the duration defined by the time search range included in the search statement, the first parameter ∈ and the second parameter MinPts involved in the density-based clustering algorithm are dynamically adjusted, which can ensure the rationality of clustering and thus improve the accuracy of the final search result.

[0353] When the search statement includes a time search range, determine the acquisition time range corresponding to each of the K (K≥1) clustering clusters; filter out the clustering clusters whose acquisition time range does not overlap with the time search range, and retain the clustering clusters whose acquisition time range overlaps with the time search range.

[0354] The overlap between the acquisition time range and the time search range can be partial overlap or full overlap. Whether it is partial overlap or full overlap between the two, it belongs to having an overlapping part.

[0355] In this way, G (G≥1) clustering clusters are selected from the K clustering clusters. The acquisition time of the earliest acquired visual media in these G clustering clusters is before the start time of the time search range, and / or, the acquisition time of the latest acquired visual media in these G clustering clusters is after the end time of the time search range. Note: There is no intersection between the G clustering clusters, and there is also no intersection between the acquisition time ranges of the G clustering clusters themselves.

[0356] The acquisition time range corresponding to the clustering cluster is from the acquisition time T min T1 of the earliest acquired visual media in the clustering cluster to the acquisition time T max of the latest acquired visual media in the clustering cluster, that is: [Tmin , T max .

[0357] Specifically, the time search range is [T start , T end , and the acquisition time range corresponding to the clustering cluster is [T min , T max . When [T min , T max and [T start , T end have an overlapping part, keep this clustering cluster; when [T min , T max and [T start , T end have no overlapping part, filter this clustering cluster. For example: the time search range is from October 1, 2022 to October 7, 2022, and the acquisition time range of the clustering cluster is from September 30, 2022 to October 1, 2022. These two ranges have an overlapping part (i.e., October 1, 2022), and this clustering cluster is retained.

[0358] Continuing with the above example, the natural scenery pictures taken by the user on September 30, 2022 and the natural scenery pictures taken by the user on October 1, 2022 are close in shooting time. Through the above clustering algorithm, they are grouped into the same clustering cluster. Assuming this clustering cluster contains only these two pictures, then the acquisition time range of this clustering cluster is: from September 30, 2022 to October 1, 2022, and it has an overlapping part with the time search range: from October 1, 2022 to October 7, 2022. Therefore, this clustering cluster will be retained, that is to say, the natural scenery pictures taken by the user on September 30, 2022 will not be filtered out.

[0359] Similarly, the natural scenery pictures taken by the user on October 8, 2022 and the natural scenery pictures taken by the user on October 7, 2022 are close in shooting time. Through the above clustering algorithm, they are grouped into the same clustering cluster. Assuming this clustering cluster contains only these two pictures, then the acquisition time range of this clustering cluster is the pictures from October 7, 2022 to October 8, 2022, and it has an overlapping part with the time search range: from October 1, 2022 to October 7, 2022. Therefore, this clustering cluster will be retained. That is to say, the natural scenery pictures taken by the user on October 8, 2022 will not be filtered out.

[0360] In this way, for the search statement "scenery pictures taken during last year's National Day holiday", the acquisition time of the earliest acquired visual media among the multiple clustering clusters obtained by screening is September 30, 2022, and the acquisition time of the latest acquired visual media is October 8, 2022.

[0361] It can be seen that using the time filtering method provided by the embodiments of the present application can ensure that a series of photos with relatively close acquisition times are presented to the user, guarantee the coherence of the search results, and improve the user's search experience.

[0362] In addition, when the search statement also includes a semantic entity related to a location, location filtering can be further performed on the G clustering clusters filtered out. Specifically, visual media in the clustering clusters with acquisition locations that do not match the semantic entity related to the location in the search statement can be filtered out. Exemplarily, the acquisition location "Prince Kung's Mansion, Xicheng District, Beijing" and the semantic entity "Sujiatuo Town, Haidian District, Beijing" can be considered to match (belonging to a municipal-level match); the acquisition location "Tsinghua University, Haidian District, Beijing" and the semantic entity "Sujiatuo Town, Haidian District, Beijing" can be considered to match (belonging to a district-level match); the acquisition location "Shanghai" and the semantic entity "Beijing" can be considered not to match.

[0363] In the above embodiments, clustering is performed first, then time filtering, and finally location filtering. Of course, in actual applications, location filtering can also be performed first, then clustering, and finally time filtering. The specific execution order of these three steps can be set according to actual needs, and the embodiments of the present application do not make specific limitations in this regard.

[0364] 5103. Filtering of character relationships.

[0365] When the search statement includes a semantic entity related to a character relationship, visual media in each clustering cluster with a character relationship attribute that does not match the semantic entity are filtered out. Exemplarily, if the search statement includes "good friends" and the name attribute of picture B is "colleague", and these two do not match, then picture B is filtered out.

[0366] 5104. Name filtering.

[0367] When the search statement includes a semantic entity related to a person's name, visual media in each clustering cluster with a name attribute that does not match the semantic entity are filtered out. Exemplarily, if the search statement includes "Zhang San" and the name attribute of picture A is "Li Si", and these two do not match, then picture A is filtered out.

[0368] It should be added that the semantic entity related to a person's name and the semantic entity related to a character relationship in the search statement can also be identified through named entity recognition technology. The name attribute and character relationship attribute of the visual media are manually added by the user for the visual media in advance.

[0369] 511. Sorting.

[0370] For the candidate set obtained in the above step 508, the visual media in the candidate set can be sorted according to the collection time sequence of the visual media in the candidate set, so as to obtain the display order of the visual media. The collection time of the visual media with a higher display order is earlier than that of the visual media with a lower display order. Subsequently, the mobile phone can display the candidate set according to the display order of the visual media in the candidate set.

[0371] For the filtered candidate set (including the above G clustering clusters) obtained in the above step 510, one of the following methods can be used for sorting:

[0372] Method 1: Sort the visual media in the G clustering clusters according to the target matching degree of the visual media in the G clustering clusters from high to low, so as to obtain the display order of the visual media in the G clustering clusters. The target matching degree can be the matching degree between the visual content of the visual media and the search statement in the above text or the comprehensive matching degree in the above text (the specific calculation method can refer to the corresponding content in the above embodiments). Subsequently, the mobile phone can display the visual media in the G clustering clusters according to the display order of the visual media in the G clustering clusters.

[0373] Method 2: Sort the G clustering clusters according to the start time sequence of the collection time ranges of the G clustering clusters, so as to obtain the display order of the G clustering clusters (that is, the inter-cluster sorting). Exemplarily, as Figure 6 shown in the interface 601, the collection time range corresponding to the clustering cluster A is from September 30, 2022 to October 1, 2022; the collection time range corresponding to the clustering cluster B is from October 3, 2022 to October 5, 2022; then, the display order of the clustering cluster A is prior to that of the clustering cluster B. For each clustering cluster, sort the visual media within the cluster according to the target matching degree of the visual media within the cluster from high to low, so as to obtain the display order between the visual media within the cluster (intra-cluster sorting); or, for each clustering cluster, sort the visual media within the cluster according to the collection time sequence of the visual media within the cluster, so as to obtain the display order of the visual media within the cluster. Subsequently, the mobile phone displays the visual media in the G clustering clusters according to the inter-cluster sorting and intra-cluster sorting.

[0374] It should be added that, in order to better protect the user privacy and security, meet the user data minimization principle, and avoid reporting user data to the cloud side as much as possible, the above entire search process is completed on the terminal side.

[0375] In addition, the present application provides an electronic device, including: a memory, a processor, and a display, wherein, the memory is used for storing a program; the processor is coupled to the memory and the display, and is used for executing the program stored in the memory to implement the above visual media search method.

[0376] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, where when the computer program is executed by a computer, one or more steps in any of the above visual media search methods can be implemented.

[0377] The computer-readable storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0378] Another embodiment of the present application also provides a computer program product containing instructions. When the computer program product is executed by a computer, one or more steps in any of the above methods can be implemented.

[0379] Among them, the electronic device, the computer-readable storage medium, and the computer program product provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here.

[0380] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces. The indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.

[0381] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they may be located in one place, or they may be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0382] In addition, the functional units in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0383] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0384] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A visual media search method applicable to an electronic device, characterized in that, it includes: displaying a first interface; the first interface includes a search box; receiving a search operation on the search statement input into the search box; determining M candidate visual media files whose visual semantic vectors match the first sentence semantic vector from multiple visual media according to the first sentence semantic vector of the search statement; the visual semantic vector of each visual media file is obtained by performing semantic understanding on the image or image frame of the visual media file using a natural picture understanding model; where M>1 and is an integer; determining N dimensions to be matched according to the semantic entities related to visual content in the search statement; N≥1 and is an integer; determining the search result of the search statement according to the matching degrees of the visual content of the M candidate visual media files on the N dimensions to be matched; displaying the search result.

2. The method according to claim 1, characterized in that, determining N dimensions to be matched according to the semantic entities related to visual content in the search statement, including: using the N semantic entities related to visual content in the search statement as the N dimensions to be matched; or using (N - 1) semantic entities related to visual content in the search statement and the search statement as the N dimensions to be matched, N>1.

3. The method according to claim 1, characterized in that, it further includes: matching the search terms in the search statement with a plurality of preset tags to determine whether the search terms belong to semantic entities related to the tags; the tags are used to describe visual content; determining the semantic entities related to the tags in the search statement as the semantic entities related to visual content in the search statement.

4. The method according to any one of claims 1 to 3, characterized in that, determining the search result of the search statement according to the matching degrees of the visual content of the M candidate visual media files on the N dimensions to be matched, including: obtaining the weights of the N dimensions to be matched; where the weight of the jth dimension to be matched is positively correlated with the degree of variation of the matching degrees of the visual content of the M candidate visual media files on the jth dimension to be matched; j is an integer, and the values of j range from 1 to N in sequence; weighted summing the matching degrees of the visual content of the ith candidate visual media file on each of the N dimensions to be matched according to the weights of the N dimensions to be matched to obtain the comprehensive matching degree of the ith candidate visual media file; i is an integer, and the values of i range from 1 to M in sequence; determining the search result of the search statement according to the comprehensive matching degrees of the M candidate visual media files.

5. The method according to claim 4, characterized in that, it further includes: determining the degree of variation of the matching degrees of the visual content of the M candidate visual media files on the jth dimension to be matched according to the information entropy of the matching degrees of the visual content of the M candidate visual media files on the jth dimension to be matched; the degree of variation is negatively correlated with the information entropy; Determine the weight of the \(j\)th dimension to be matched according to the degree of variation of the matching degree of the visual content of the \(M\) candidate visual media files on the \(j\)th dimension to be matched.

6. The method according to claim 4, wherein, further comprising: Determine the initial matching degree of the \(M\) candidate visual media on the \(j\)th dimension to be matched according to the vector similarity between the visual semantic vector of the \(M\) candidate visual media files and the semantic vector to be matched of the \(j\)th dimension to be matched; Perform normalization processing on the initial matching degree of the \(M\) candidate visual media on the \(j\)th dimension to be matched to obtain the matching degree of the \(M\) candidate visual media on the \(j\)th dimension to be matched.

7. The method according to claim 4, wherein, Determine the search result of the search statement according to the comprehensive matching degree of the \(M\) candidate visual media files, including: Determine multiple first visual media files whose comprehensive matching degree meets the preset requirements from the \(M\) candidate visual media files; Determine the search result of the search statement according to the multiple first visual media files.

8. The method according to claim 7, wherein, Determine the search result of the search statement according to multiple first visual media files, including: If the search statement includes a semantic entity related to time, filter the multiple first visual media files according to the semantic entity related to time and the acquisition time attribute of the multiple first visual media files; If the search statement includes a semantic entity related to location, filter the multiple first visual media files according to the semantic entity related to location and the acquisition location attribute of the multiple first visual media files; If the search statement includes a semantic entity related to the relationship between people, filter the multiple first visual media files according to the semantic entity related to the relationship between people and the relationship between people attribute of the multiple first visual media files; and / or If the search statement includes a semantic entity related to a person's name, filter the multiple first visual media files according to the semantic entity related to the person's name and the person's name attribute of the multiple first visual media files.

9. An electronic device, wherein, comprising: A memory, a processor, and a display, wherein, The memory is used to store programs; The processor is coupled to the memory and the display, and is used to execute the program stored in the memory to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, wherein, The computer program, when executed by a computer, is capable of implementing the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Abstract generating method of extensible markup language (XML) keyword search

    CN102004802A

  • Enterprise name retrieval method, enterprise name retrieval device and terminal equipment

    CN112597208A

  • Visual media personalized search method and device

    CN113641857A

  • Video searching method and device, electronic equipment and storage medium

    CN115017361A

  • Similar image retrieval

    US20150169740A1

Cited By

  • Information display method and electronic device

    WO2026123859A1