Visual media search method, device, and storage medium
By improving the Porta count method and multi-dimensional fusion sorting, and combining user click counts and visual semantic similarity, the accuracy problem of image library applications when processing complex search queries has been solved, achieving more accurate visual media search and improving user experience.
Patent Information
- Application Number
- CN202311571222.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2043-11-21
AI Technical Summary
Existing image library applications cannot accurately understand users' search intent when processing complex search queries, resulting in inaccurate search results and failing to meet users' precise search needs.
An improved Porta count method is used for multi-dimensional fusion ranking. It combines user click count, visual semantic vector of visual media file with sentence-to-sentence similarity of search terms, as well as visual content-irrelevant attributes. The scores of visual media files are calculated by using tags generated by a natural image understanding model, thereby improving the accuracy of search results.
By integrating and sorting data across multiple dimensions, the accuracy of visual media search and user experience are improved. It can handle complex search queries, provide more accurate search results, reduce computational load, and lower power consumption.
Smart Images

Figure CN120067371B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of terminal, and particularly to a visual media search method, device and storage medium. BACKGROUND
[0002] With the popularity of intelligent terminals, more and more users use intelligent terminals, such as mobile phones, to take photos and videos, and store the taken photos and videos in the gallery of an electronic device, so as to record the details of life. In addition, users can also download pictures, take screenshots of the interface of the mobile phone, and store the downloaded pictures and screenshots in the gallery of the electronic device.
[0003] In order to facilitate users to manage and view pictures in the terminal, the gallery application or other similar applications of the terminal device are configured with a picture management function and a picture search function. For example, the gallery application in the terminal can classify the pictures in the terminal according to the time and location of taking the pictures, and generate corresponding albums, and the user can view related pictures by searching for time and location information. SUMMARY
[0004] Aspects of the present application provide a visual media search method, device and storage medium, which can improve the rationality of score statistics of each visual media, and further improve the user search experience.
[0005] In a first aspect, a visual media search method is provided, which is suitable for an electronic device and includes:
[0006] displaying a first interface; the first interface includes a search box;
[0007] receiving a search operation of inputting a search statement into the search box;
[0008] determining W matching dimensions according to the search statement; W is an integer greater than 1;
[0009] obtaining the matching degree and the matching degree ranking of Y visual media files in the wth matching dimension; Y is an integer greater than 1; w is an integer and takes a value from 1 to W;
[0010] correcting the total number of times that the yth visual media file appears in the uth according to the matching degree of the yth visual media file in each matching dimension in the W matching dimensions, to obtain the corrected total number of times that the yth visual media file appears in the uth; wherein u is an integer and takes a value from 1 to Y; y is an integer and takes a value from 1 to Y;
[0011] determining the score of the yth visual media according to the corrected total number of times that the yth visual media file appears in each ranking;
[0012] determine a search result of the search sentence according to the scores of the Y visual media files;
[0013] display the search result.
[0014] Specifically, the Y visual media files can be ranked according to the scores of the Y visual media files to obtain the search result of the search sentence, or the Y visual media files can be filtered according to the scores of the Y visual media files to obtain the search result of the search sentence.
[0015] That is, in the visual media search scenario, when the scores of the visual media files are counted based on the ranking of the matching degrees of the visual media files in each to-be-matched dimension, the matching degree of the visual media file in each to-be-matched dimension is additionally introduced to correct the total number of times that the visual media file appears at each rank, which helps to improve the rationality of the score counting and further improve the user search experience.
[0016] It can be understood that the scheme is a comprehensive ranking of the visual media files based on the improved borda counting method in multiple dimensions.
[0017] In one possible implementation, the total number of times that the yth visual media file appears at the u-th rank is corrected according to the matching degree of the yth visual media file in each to-be-matched dimension in the W to-be-matched dimensions, to obtain a corrected total number of times that the yth visual media file appears at the u-th rank, including:
[0018] determining a first weight of the w-th to-be-matched dimension for the yth visual media file according to the matching degree of the yth visual media file in the w-th to-be-matched dimension;
[0019] performing weighted summation on the number of times that the yth visual media file appears at the u-th rank in each to-be-matched dimension in the W to-be-matched dimensions according to the first weights of the W to-be-matched dimensions for the yth visual media file, to obtain the corrected total number of times that the yth visual media file appears at the u-th rank.
[0020] In one possible implementation, the first weight of the w-th to-be-matched dimension for the yth visual media file is determined according to the matching degree of the yth visual media file in the w-th to-be-matched dimension, including:
[0021] obtaining a preset weight of the w-th to-be-matched dimension;
[0022] According to a preset weight of the wth dimension to be matched, the matching degree of the yth visual media file on the wth dimension to be matched is corrected to obtain a corrected matching degree of the yth visual media file on the wth dimension to be matched.
[0023] The corrected matching degree of the yth visual media file on the wth dimension to be matched is determined as the first weight of the wth dimension to be matched for the yth visual media file.
[0024] That is, a weight configuration channel is reserved for each dimension in the comprehensive ranking, and the preset weight of each dimension can be set according to prior knowledge.
[0025] In one possible implementation, the determination manner of the dimensions to be matched includes multiple manners in the following manners:
[0026] The search statement is determined as a dimension to be matched;
[0027] A semantic subject irrelevant to visual content in the search statement is determined as a dimension to be matched;
[0028] The number of user clicks is determined as a dimension to be matched.
[0029] The number of user clicks can represent the personalized preference of the user. When the visual media is comprehensively ranked in multiple dimensions, the personalized preference of the user and other information are considered, which can improve the rationality of score statistics and further improve the user search experience.
[0030] In one possible implementation, the W dimensions to be matched include the number of user clicks.
[0031] According to the number of times that the yth visual media file is clicked by the user in history, a matching degree of the yth visual media file on the number of user clicks is determined.
[0032] In one possible implementation, the W dimensions to be matched include the search statement.
[0033] According to a vector similarity between a visual semantic vector of the yth visual media file and a first sentence semantic vector of the search statement, a matching degree of the yth visual media file on the search statement is determined.
[0034] The matching degree of the yth visual media file on the search statement represents the matching degree of the visual content of the visual media file and the entire search statement. That is, when the comprehensive ranking is performed, the matching degree of the visual content of the visual media file and the entire search statement is considered.
[0035] In a possible implementation, the W dimensions to be matched include a semantic subject irrelevant to visual content in the search statement.
[0036] According to the matching of the semantic subject irrelevant to visual content in the search statement according to the attribute of the yth visual media file, a matching degree of the yth visual media file on the semantic subject irrelevant to visual content in the search statement is determined.
[0037] That is, when the comprehensive ranking is performed, the matching degree of the attribute of the visual media file such as time and place with the search statement is considered.
[0038] In a possible implementation, before the total number of uth occurrences of the yth visual media file is corrected according to the matching degrees of the yth visual media file on each of the W dimensions to be matched, the method further includes:
[0039] The matching degrees of the Y visual media files on the wth dimension to be matched are processed in percentage.
[0040] In this solution, the matching degree on the dimension of the number of user clicks is the actual number of user clicks, and in order to ensure the effectiveness of subsequent calculation, the matching degree on the dimension of the number of user clicks needs to be processed in percentage. Specifically, an optimal matching degree on the dimension of the number of user clicks can be set, and then the matching degree on the dimension of the number of user clicks is processed in percentage based on the optimal matching degree.
[0041] In a possible implementation, the method further includes:
[0042] The search statement is subjected to semantic understanding to obtain a first sentence semantic vector;
[0043] Visual semantic vectors of a plurality of visual media files are obtained; the visual semantic vector of each visual media file is obtained by subjecting an image or an image frame of the visual media file to semantic understanding by using a natural picture understanding model;
[0044] From the plurality of visual media files, a plurality of first visual media files having visual semantic vectors matching the first sentence semantic vector are determined;
[0045] The Y visual media files are determined according to the plurality of first visual media files.
[0046] In the solution, the Y visual media files are recalled based on matching the visual semantic vectors of the visual media files with the sentence semantic vector of the search statement. That is, the visual content of the Y visual media files is semantically matched with the whole search statement, which is consistent with the visual attention of the user. Subsequently, only the visual media files consistent with the visual attention of the user need to be comprehensively sorted, so that the search accuracy can be ensured, the calculation amount involved in the comprehensive sorting can be reduced, the search delay can be shortened, and the power consumption can be reduced.
[0047] In a second aspect, the present application provides an electronic device, comprising: a memory, a processor and a display, wherein,
[0048] The memory is configured to store a program.
[0049] The display is configured to display a search page.
[0050] The processor is coupled with the memory and the display, and is configured to execute the program stored in the memory to implement the method.
[0051] In a third aspect, the present application provides a computer readable storage medium storing a computer program, wherein the computer program is executed by a computer to implement the method. BRIEF DESCRIPTION OF DRAWINGS
[0052] The accompanying drawings, which are included to provide a further understanding of the present application, constitute a part of the specification and illustrate the illustrative embodiments of the present application and the description thereof, and are used to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:
[0053] Figure 1A A set of interface diagrams of a search interface of a mobile phone entering a gallery application provided by an embodiment of the present application;
[0054] Figure 1B A search interface diagram after a search history is cleared provided by an embodiment of the present application;
[0055] Figure 1C A set of interface diagrams involved in searching in a gallery application provided by an embodiment of the present application;
[0056] Figure 1D A search failure interface diagram provided by an embodiment of the present application;
[0057] Figure 2A A structural diagram of an electronic device provided by another embodiment of the present application;
[0058] Figure 2B A software structure block diagram of an electronic device provided by another embodiment of the present application;
[0059] Figure 3A Another set of interface diagrams related to searching in a gallery application provided by an embodiment of the present application;
[0060] Figure 3B A set of interface diagrams related to searching in a negative one screen provided by an embodiment of the present application;
[0061] Figure 3C Search result interface diagram one provided by an embodiment of the present application;
[0062] Figure 3D Search result interface diagram two provided by an embodiment of the present application;
[0063] Figure 3E Search result interface diagram three provided by an embodiment of the present application;
[0064] Figure 3F Search result interface diagram four provided by an embodiment of the present application Figure Four ;
[0065] Figure 4 Interaction diagram of a visual media search method provided by an embodiment of the present application;
[0066] Figure 5 Flowchart of a visual media search method provided by an embodiment of the present application;
[0067] Figure 6 Search result interface diagram five provided by an embodiment of the present application Figure Five ;
[0068] Figure 7 Flowchart of a visual media search method provided by an embodiment of the present application. DETAILED DESCRIPTION
[0069] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B; in this document, "and / or" is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.
[0070] Hereinafter, the terms "first", "second", "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second", "third" can explicitly or implicitly include one or more of the features.
[0071] Firstly, the glossary involved in the embodiments of the present application is explained. It can be understood that the explanation is for a clearer understanding of the embodiments of the present application, and does not necessarily constitute a limitation on the embodiments of the present application.
[0072] Visual media: refers to pictures or videos.
[0073] Semantic subject: the named entity recognition technology can recognize the text, recognize the entities with specific meaning such as names (PER) and places (LOC) in the text. In the present scheme, the identified entities with specific meaning are called semantic subjects.
[0074] Visual content related and visual content independent: visual content refers to the content of the objects displayed by visual media and their mutual relationship. Computer vision can enable computers to have similar visual capabilities to humans, including perception, understanding, analysis and interpretation of visual content. At present, the generative pre-trained transformer model 4 (GPT-4) can support the input of images to the model, and output the human natural language describing the important information in the image.
[0075] In the image search context of the present scheme, the data that the visual media file needs to obtain through the natural picture understanding of the model is called "visual content related". In the following, the visual content is the data that needs to be obtained through the natural picture understanding model. The data related to the visual media file and can be obtained without the picture understanding ability of the model is called "visual content independent" in the present scheme, such as the shooting location, shooting time, name, file attribute, etc. which can be obtained and saved by the terminal device when collecting the visual media file.
[0076] For example, "photos taken in Beijing this year" in "this year" (shooting time), "Beijing" (shooting location), "photos" (file attribute) are data that can be obtained and saved by the terminal device when collecting the visual media file, so "this year", "Beijing", "photos" are independent of visual content; "sky" in "sky in Beijing this year" needs to be understood through the picture understanding ability of the model to the image or image frame of the visual media file, so "sky" is visual content related.
[0077] Text semantic vector: can be obtained by sending the text into a text encoder, which can represent the semantic feature vector of the entire sentence. The text encoder can use the commonly used Transformer model in natural language processing (NLP), and the present scheme does not limit it. In the present scheme, the text semantic vector obtained for the sentence is called a sentence semantic vector, the text semantic vector obtained for the semantic subject in the sentence is called a subject semantic vector, and the text semantic vector obtained for the label is called a label semantic vector.
[0078] Visual semantic vector: can be obtained by sending the image or image frame of the visual media file into an image encoder. The commonly used CNN (Convolutional Neural Network) model or VIT (Vision Transformer) model can be used, and the present scheme does not limit it.
[0079] Density-based clustering algorithm: is based on a set of neighborhoods to describe the closeness of the sample set, and the first parameter ∈ and the second parameter MinPts are used to describe the closeness of the sample distribution of the neighborhood. The first parameter ∈ is used to describe the neighborhood radius of the data point; the second parameter MinPts is used to describe the minimum number of data points in the neighborhood of the data point. Its representative algorithm is DBSCAN (Density-Based Spatial Clustering of Application with Noise, Density-Based Spatial Clustering of Application with Noise); DBSCAN algorithm is a representative density-based clustering algorithm, which can divide regions with high enough density into clusters, and can find clusters of arbitrary shape in a noisy spatial database;
[0080] Vector similarity: used to describe the similarity between two vectors (for example: between the sentence semantic vector and the visual semantic vector). In the embodiments of the present application, the similarity between the sentence semantic vector and the visual semantic vector can be compared to determine the visual media that matches the search statement. Generally, the vector similarity can be calculated by the cosine similarity formula, and of course, it can also be calculated by other means.
[0081] In the prior art, a mobile phone manages the visual media files (hereinafter referred to as visual media) of the user's pictures, videos, etc. through a gallery application. Taking a mobile phone to take a photo as an example, after the mobile phone takes a photo, the gallery application can obtain and save the shooting location, shooting time, photo name and other attributes unrelated to the visual content of the photo.
[0082] In practical applications, photos can be input into a natural image understanding model while the phone is charging and the screen is off. The model generates and saves tags for the photos, which can be seen as attributes related to the visual content of the photo. Tags could be "sky," "cat," "dog," etc. The gallery app can then index the photos based on these attributes. Once the index is established, the gallery app can provide users with corresponding search services. Specifically, users can search for images or videos in the gallery app by entering keywords. For example, users can enter keywords such as "Beijing," "sky," or "National Day" in the search box provided by the gallery app. The gallery app will then match the entered keywords with the index of images, videos, and other visual media within the app to obtain search results.
[0083] The following description, in conjunction with the accompanying drawings, illustrates the interface involved in the search process of a gallery application in the prior art:
[0084] like Figure 1A As shown in (a), the mobile phone can display a main interface 101, which can also be referred to as the desktop. The main interface 101 may include an icon 102 for the gallery application. The mobile phone can receive a user's click on icon 102; in response to this action, the mobile phone can launch the gallery application and display... Figure 1A The interface 103 shown in (b) can be a photo album interface. It should be noted that, in response to the user clicking icon 102, the phone can launch the gallery application and display the photo gallery interface. The photo gallery interface includes thumbnails of photos (i.e., images) in the gallery or a large image of a specific photo. Within the photo gallery interface, in response to the user's operation of the "Album" control, the aforementioned album interface 103 is displayed.
[0085] like Figure 1A As shown in (b), interface 103 includes multiple albums, including "All Photos" album containing 2023 photos, "Camera" album containing 1502 photos and videos, "Screenshot and Screen Recording" album containing 102 photos and videos, "My Favorites" album containing 48 photos and videos, "One Record, Multiple Views" album containing 34 photos and videos, "Video Editing" album containing 65 videos, "Custom Albums" album containing 57 photos and videos, and "Shared Albums" album containing 100 photos and videos.
[0086] like Figure 1A As shown in (b), the interface 103 may include a search box 104. The mobile phone can receive the user's click on the search box 104, and in response to this operation, the mobile phone can display as shown in Figure (b). Figure 1AThe interface 105 shown in (c) of FIG. 1 can be referred to as a search interface. In the interface 105, the phone can show the user the classification information of the photos. For example, in the interface 105, the phone classifies the local photos according to time, person, and thing, etc. For example, in the time dimension, the phone classifies the local photos into three time periods of "this month", "last month", and "this year", where the "this month" album includes the photos or videos taken by the phone in this month, the "last month" album includes the photos or videos taken by the phone in last month, and the "this year" album includes the photos or videos taken by the phone in this year. In the person dimension, the phone classifies the local photos according to different persons, for example, the four different persons in the interface 105. In the thing dimension, the phone classifies and shows the local photos according to "landscape", "animal", "document", and "building". It should be noted that the classification dimensions described above can also be other dimensions, which are not specifically limited herein. In the interface 105, the user can see the classification information without inputting a keyword.
[0087] Optionally, the interface 105 can also include a search history 107 and a "clear" 108 option. The search history includes the keywords that the user has inputted before, for example, "flowers", "coffee", "cat", etc. The phone can receive the user's operation of clicking the "clear" 108, and in response to the operation, the phone can clear the search history. After the phone clears the search history, the search interface 105 no longer displays the keywords that the user has inputted before. For example, in response to the user's operation of clicking the "clear" 108, as shown in (b) of FIG. 1, the search interface 105 no longer displays the search history 107 and the "clear" 108 option, and the content displayed below is moved up. Figure 1B
[0088] In response to the user's operation of inputting the keyword "sky" in the interface 105, the phone displays the interface 109 shown in (a) of FIG. 1. As shown in (a) of FIG. 1, there are 100 photos related to "sky" and 32 photos containing the word "sky". Figure 1C Figure 1C The 100 photos related to "sky" can be recalled because the tags of the 100 photos match "sky"; the 32 photos containing the word "sky" can be recalled because the OCR (Optical Character Recognition) technology identifies that the 32 photos contain the word "sky". In actual applications, the phone can also associate the keyword inputted by the user to obtain an associated keyword, and perform a search based on the associated keyword.
[0089] The interface 109 also displays some search results related to the keyword "sky" and a "More" option 110 corresponding to the search results of the keyword "sky". The mobile phone receives a click operation by the user on the "More" option 110 and displays an interface 111 as shown in Figure 1C (b) shown in the figure. Among them, the interface 111 is used to display photos and videos in the search results of the keyword "sky". Optionally, the photos and videos can be classified and displayed according to time. In addition, the interface 111 also includes a return key 112 and a title 113. In response to the user's operation on the return key 112, the mobile phone can redisplay the interface 109. The title 113 may include the keyword "sky".
[0090] That is to say, in the existing gallery applications, when the user enters simple keywords in the search box, such as: sky, Beijing, National Day, etc., corresponding search results can be obtained. However, because the number of photo attributes is simple and limited, and the mobile phone's ability to understand and associate search statements is also limited. If the user enters a more complex search statement in the search box, and the keywords in the search statement cannot match the attributes of the picture or the text in the picture, no photos can be found. That is to say, the existing gallery applications do not support the search function based on complex search statements. As Figure 1D shown, when the user enters a more complex search statement "warming oneself by the fire and boiling tea" in the search box of the interface 114, the mobile phone cannot understand the associated words of "warming oneself by the fire and boiling tea", and since the photos do not have attributes that can match "warming oneself by the fire and boiling tea" or its associated words, the search result shows "no pictures".
[0091] However, in practical applications, users have a strong demand for the function of searching for pictures based on complex search statements. This is because users can describe the pictures or videos they want more comprehensively through complex search statements, thereby achieving accurate search. To meet this demand of users, the embodiments of the present application provide a visual media search method. This method can be applied to electronic devices, and the electronic devices can be terminal devices such as mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc.
[0092] Exemplarily, Figure 2AA structural diagram of the electronic device 200 is shown. The electronic device 100 can include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headset jack 270D, a sensor module 280, a key 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.
[0093] The sensor module 280 can include a pressure sensor 280A, a gyro sensor 280B, a barometric sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.
[0094] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 can include more or fewer components than shown, or combine certain components, or split certain components, or different component arrangements. The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0095] The processor 210 can include one or more processing units, for example: the processor 210 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices, or can be integrated in one or more processors.
[0096] The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.
[0097] The processor 210 can also include memory that stores instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. The memory can hold instructions or data that the processor 210 has recently used or has frequently used. If the processor 210 needs to use the instructions or data again, it can call them directly from the memory. This avoids repeated access and reduces the latency of the processor 210, thus improving the efficiency of the system.
[0098] In some embodiments, the processor 210 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0099] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a limitation on the structure of the electronic device 200. In some other embodiments of the present application, the electronic device 200 can also use different interface connection modes or a combination of multiple interface connection modes in the above embodiments.
[0100] The electronic device 200 realizes the display function through the GPU (Graphics Processing Unit, graphics processor), the display screen 294, and the application processor, etc. The GPU is connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 can include one or more GPUs that execute program instructions to generate or change display information.
[0101] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diodes (QLED), etc. In some embodiments, the electronic device 100 can include 1 or N display screens 294, N being a positive integer greater than 1.
[0102] The electronic device 200 can implement the photographing function through the ISP, the camera 293, the video codec, the GPU, the display screen 294, and the application processor, etc.
[0103] The camera 293 is used to capture still images or videos. An object generates an optical image through a lens and projects it to a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB, YUV, etc. In some embodiments, the electronic device 100 can include 1 or N cameras 293, N being a positive integer greater than 1.
[0104] The video codec is used to compress or decompress digital videos. The electronic device 200 can support one or more video codecs. In this way, the electronic device 200 can play or record videos in multiple encoding formats, such as moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.
[0105] The NPU is a neural-network (NN) computing processor. By drawing on the structure of a biological neural network, for example, by drawing on the transmission mode between human brain neurons, the NPU can quickly process input information and can also constantly self-learn. Through the NPU, intelligent cognition and other applications of the electronic device 200 can be implemented, for example: image recognition, face recognition, speech recognition, text understanding, and the like.
[0106] The external memory interface 220 can be used to connect an external memory card, for example, a Micro SD card, to realize the expansion of the storage capability of the electronic device 200. The external memory card communicates with the processor 210 through the external memory interface 220 to realize a data storage function. For example, music, video, and the like are saved in the external memory card.
[0107] The internal memory 221 can be used to store computer executable program codes, which include instructions. The internal memory 221 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, and the like), and the like. The data storage area can store data (such as audio data, a phone book, and the like) created during the use of the electronic device 200, and the like. In addition, the internal memory 221 can include a high-speed random access memory and can also include a non-volatile memory, for example, at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), and the like. The processor 210 executes various function applications and data processing of the electronic device 200 by running instructions stored in the internal memory 221 and / or instructions stored in a memory disposed in the processor.
[0108] Figure 2B FIG. 2 is a software structure block diagram of the electronic device 200 according to an embodiment of the present disclosure. The software system of the electronic device 200 can adopt a layered architecture, which divides the software into several layers, each of which has a clear role and division of labor. Layers communicate with each other through software interfaces. Taking the Android system as an example, in some embodiments, the Android system can be divided into four layers, from top to bottom, an application layer, an application framework layer, an Android runtime and a system library, and a kernel layer.
[0109] As shown in FIG. 2, the application layer can include a gallery service module, a search module, a multi-modal understanding module, a natural language understanding module, a camera application, and the like. Figure 2B
[0110] The application framework layer provides an application programming interface (API) and a programming framework for applications of the application layer. The application framework layer includes some pre-defined functions.
[0111] The system library can include a plurality of functional modules. For example, a surface manager, media libraries, a three-dimensional graphics processing library (e.g., OpenGL ES), a 2D graphics engine (e.g., SGL), and the like. Among them, the media libraries support playback and recording of a plurality of commonly used audio, video formats, and static image files, and the like. The media libraries can support a plurality of audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, and the like.
[0112] The kernel layer is a layer between hardware and software.
[0113] The following describes, by way of example, a workflow of the software and hardware of the electronic device 200 in a capture photographing scenario.
[0114] When the touch sensor 280K receives a touch operation, a corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including touch coordinates, a timestamp of the touch operation, and the like). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer, and identifies a control corresponding to the input event. Taking an example in which the touch operation is a touch single-click operation and the control corresponding to the single-click operation is a control of a camera application icon, the camera application invokes an interface of the application framework layer, starts the camera application, and then starts the camera driver by invoking the kernel layer, and captures a still image or a video by the camera 293.
[0115] The following describes, by way of example, a workflow of the software and hardware of the electronic device 200 in a capture photographing scenario.
[0116] The following describes, by way of example, a workflow of the software and hardware of the electronic device 200 in a capture photographing scenario.
[0117] As shown in interface 301 in (a) of FIG. 3, the interface 301 displays a search history 303 and a "clear" option 304. The search history includes search statements that have been input by the user, such as "watching sunrise on the top of a mountain" and "the Great Wall photographed in Beijing this year". The interface 301 can display other content as described above with reference to the interface 305, which is not repeated here. The interface 301 can be referred to as a search interface. The mobile phone can respond to a user input of a search operation on the search interface 301. Figure 3A As shown in interface 301 in (a) of FIG. 3, the interface 301 displays a search history 303 and a "clear" option 304. The search history includes search statements that have been input by the user, such as "watching sunrise on the top of a mountain" and "the Great Wall photographed in Beijing this year". The interface 301 can display other content as described above with reference to the interface 305, which is not repeated here. The interface 301 can be referred to as a search interface. The mobile phone can respond to a user input of a search operation on the search interface 301.Figure 1A Clicking on the search box 104 in the interface 103 shown in (b) in [reference], the interface 301 is displayed.
[0118] As Figure 3A In the interface 305 (i.e., the search result interface) shown in (b) in [reference], the user enters the search statement "warming the stove and boiling tea" in the search box 306 of the interface 305, and the mobile phone searches for 239 pictures. The mobile phone displays some search results (e.g., thumbnails of 8 pictures) of the search statement "warming the stove and boiling tea" and the "More" option 307 corresponding to the search results of the search statement "warming the stove and boiling tea" on the interface 305. In response to the user's operation on the "More" option 307, the mobile phone displays the interface 308 shown in (c) in [reference]. The interface 308 can display the pictures in the search results of the search statement "warming the stove and boiling tea" in descending order according to the matching degree of the visual content of the pictures and the search statement "warming the stove and boiling tea" (specifically, it can be the similarity between the visual semantic vector of the picture and the sentence semantic vector of the search statement). The user can also perform an upward sliding operation on the interface 308 to view the pictures that have not been displayed. Figure 3A
[0119] Figure 3B Currently, the mobile phone also has a negative first screen, a pull-down search interface, etc. It can be understood that the negative first screen can be the leftmost split screen of the electronic device, which is used to provide functions such as search and quick services for users. Among them, the negative first screen can also be used to display notification messages to be pushed to users, such as application messages subscribed by users, real-time hot search messages, segment selection, itinerary information, etc. The pull-down search interface is an interface displayed in response to the user's pull-down operation on the main interface, and this interface is used to provide functions such as search and application suggestions for users, and this interface and The interface
[0122] The mobile phone receives a click operation of the user on the search box 311 on the negative one screen 310, and displays an interface 315 as shown in (c) of Figure 3B The interface 315 can include a search box 316, application suggestions, and a search history 217. The application suggestions include icons of applications that are suggested to be used. The interface 215 can also include the search history 217 and a corresponding "clear" option 318. In response to a triggering operation of the user on the "clear" option 318, the search history 317 and the "clear" option 318 are no longer displayed on the interface 315. In addition, a hot news title, for example, "Tianjin Marathon", can be displayed in the search box 316.
[0123] As shown in (d) of Figure 3B , the search box of the interface 319 displays a search sentence "Great Wall shot in Beijing during National Day" input by the user; the interface 319 also displays a preview area 322 of the search result of the gallery application for the search sentence "Great Wall shot in Beijing during National Day" and a corresponding "search in application" option 323 of the gallery application. In response to a triggering operation of the user on the preview area 322, a photo detail interface provided by the gallery application is entered for the user to page through the search result of the search sentence "Great Wall shot in Beijing during National Day". In response to a triggering operation of the user on the "search in application" option 323, the mobile phone displays an interface 324 provided by the gallery application as shown in (e) of Figure 3B , which displays part of the search result of the search sentence "Great Wall shot in Beijing during National Day" and a "more" option corresponding to the search result of the search sentence "Great Wall shot in Beijing during National Day". In response to a triggering operation of the user on the "more" option, the mobile phone can display a search result detail interface, which displays the pictures in the search result of the search sentence "Great Wall shot in Beijing during National Day". The interface 319 can also display an online search option 321. In response to a triggering operation of the user on the online search option 321, the mobile phone displays a search web page and displays the online search result in the search web page.
[0124] As shown in (f) of Figure 3CAs shown in interface 325, when a user enters the search term "Great Wall photographed during last year's National Day" in the search box, the phone retrieves 419 images. The phone displays thumbnails of each image or video in the search results for "Great Wall photographed during last year's National Day" on interface 325. Thumbnail A corresponds to an image or video taken at 23:22 on September 30, 2022; thumbnail B corresponds to an image or video taken at 00:19 on October 1, 2022. The search range for "National Day" is from 00:00 on October 1, 2022 (the first time point) to 00:00 on October 8, 2022 (the second time point). The image or video corresponding to thumbnail A was taken before 00:00 on October 1, 2022. The images or videos corresponding to thumbnail B were taken between 00:19 on October 1, 2022 and 00:00 on October 8, 2022. The images or videos corresponding to thumbnails A and B were taken relatively close in time.
[0125] like Figure 3D Interface 326, for the search query "sky photographed during last year's National Day holiday," displays thumbnail C, whose corresponding image or video was taken at 22:19 on October 7, 2022; and image D, whose corresponding image or video was taken at 01:24 on October 8, 2022. The search range for "National Day" is from 00:00 on October 1, 2022 (the first time point) to 00:00 on October 8, 2022 (the second time point). The image or video corresponding to thumbnail D was taken after 00:00 on October 8, 2022. The image or video corresponding to thumbnail C was taken between 00:00 on October 1, 2022 and 00:00 on October 8, 2022, when the time was 00:19 on October 1, 2022. Among them, the images or videos corresponding to thumbnails C and D were captured at relatively close times.
[0126] In practical applications, search queries can include not only a time-based search range (e.g., National Day) but also a primary keyword (e.g., the Great Wall or the sky). Therefore, for each search query, the visual media files corresponding to the search results are matched with the primary keyword.
[0127] When the first keyword belongs to a semantic subject unrelated to the visual content (e.g., location), the first keyword can be matched with the attributes of the visual media to determine the visual media file that matches the first keyword.
[0128] When the first keyword belongs to a semantic subject related to visual content (for example, the Great Wall or the sky), the first keyword is matched with the tags of the visual media files to determine the visual media files in which the visual content matches the first keyword; or the text semantic vector of the first keyword is matched with the visual semantic vector of the visual media files to determine the visual media files in which the visual content matches the first keyword (including the first, second, and third visual media files described above).
[0129] The thumbnail is an equal-ratio thumbnail of the corresponding visual media file or an equal-ratio thumbnail of any image frame of the corresponding visual media file.
[0130] As shown in an interface 327, for the search statement "sky shot last year during the National Day", two thumbnails E and F are displayed in order of the matching degrees of the visual content of the respective visual media files to the first keyword "sky" from large to small. The matching degree of the visual content of the visual media file corresponding to the thumbnail E to "sky" is 0.87, and the matching degree of the visual content of the visual media file corresponding to the thumbnail F to "sky" is 0.76. Figure 3E As shown in an interface 328, for the search statement "photos shot in Nankai District, Tianjin last year during the National Day", two thumbnails J and H are displayed in order of the matching degrees of the location attributes of the respective visual media files to the first keyword "Nankai District, Tianjin" from large to small. The matching degree of the location attribute of the visual media file corresponding to the thumbnail J to "Nankai District, Tianjin" is 0.87, and the matching degree of the location attribute of the visual media file corresponding to the thumbnail H to "Nankai District, Tianjin" is 0.8.
[0131] Figure 3F An interaction diagram of the visual media search method provided by the embodiments of the present application is shown in FIG. 4.
[0132] As shown in FIG. 4, the mobile phone is provided with a gallery service module (i.e., a gallery application) 41, a search module 42, a multi-modal understanding module 43, and a natural language understanding module 44. Figure 4 Figure 4 As shown in FIG. 5, the visual media search method provided by the embodiments of the present application can be divided into two stages: an index construction stage and a search stage.
[0133] As shown in FIG. 5, the visual media search method provided by the embodiments of the present application can be divided into two stages: an index construction stage and a search stage. Figure 4 In the index construction stage, the following steps are included:
[0134] S401, add and / or modify visual media and attributes thereof.
[0135]
[0136] In S401, for the newly added visual media, the gallery application can automatically generate attributes irrelevant to the visual content, which can include but are not limited to: collection location, collection time, and visual media name. Taking a video or picture as an example, the collection location refers to the shooting location, and the collection time refers to the shooting time; taking a screenshot as an example, the collection location refers to the screenshot location, and the collection time refers to the screenshot time; taking a downloaded video or picture as an example, the collection location refers to the download location, and the collection time refers to the download time.
[0137] The user can add visual media by shooting, downloading, taking screenshots, and the like. In addition, the user can also modify existing visual media. The modification includes but is not limited to beautifying, customizing the name, adding a watermark, and the like.
[0138] In S402, the visual media and its attributes are stored.
[0139] The gallery service module 41 can store the visual media and its attributes locally on the phone in response to the above-mentioned adding or modifying operation. In actual application, under the authorization of the user, the phone can store the locally stored visual media and its attributes to the cloud to reduce the storage pressure on the phone locally.
[0140] In S403, a request is made for visual semantic understanding of the visual media.
[0141] Since visual semantic understanding requires a large amount of computing resources, in order not to affect the user's use, the above-mentioned step S403 can be performed when the phone is in a charging and screen-off state.
[0142] The gallery service module 41 can request the multimodal understanding module 43 to perform visual semantic understanding on the visual media obtained by adding or modifying, to obtain the visual semantic vector of the visual media.
[0143] The multimodal understanding module 43 can perform visual semantic understanding on the visual media based on a multimodal model to obtain the visual semantic vector of the visual media.
[0144] The multimodal model is used not only for visual semantic understanding of the visual media to obtain the visual semantic vector of the visual media, but also for semantic understanding of the search statement to obtain the sentence semantic vector of the search statement, semantic understanding of the rewritten search statement in the following text to obtain the sentence semantic vector of the rewritten search statement, semantic understanding of the semantic subject in the search statement to obtain the subject semantic vector of the semantic subject, and semantic understanding of the label of the visual media to obtain the label semantic vector of the label. The multimodal model can be obtained by training according to training samples.
[0145] For example, a multimodal model can specifically be CLIP (Contrastive Language-Image Pre-training, a pre-trained model based on contrastive text-image pairs). The CLIP model can map visual media and text (i.e., search statements, rewritten search statements, semantic subjects, and tags) to a unified vector space to understand the relationships between different modal resources in both text and visual terms, thereby enabling image retrieval. Specifically, in this embodiment, the CLIP model can be used to match visual media files with text.
[0146] Multimodal models can map visual media and text into vectors of the same dimension. That is, the dimension of the visual semantic vector of visual media is the same as the dimension of the semantic vector of text (e.g., the sentence semantic vector of a search query). A multimodal model includes, as mentioned above, an image encoder and a text encoder.
[0147] S404. Returns the visual semantic vector of the visual media.
[0148] The multimodal understanding module 44 returns the visual semantic vector of the visual media to the image library service module 41.
[0149] S405, Store the visual semantic vector of the visual media.
[0150] The gallery service module 41 can store visual semantic vectors of visual media locally.
[0151] S406. Send the attribute information of the visual media and its visual semantic vector.
[0152] For example, the image library service module 41 can store the visual semantic vectors of the visual media returned by the multimodal understanding module 44, and then send the attributes of the visual media and their visual semantic vectors to the search module 42 in batches, so that the search module 42 can build an index of the visual media.
[0153] S407, Building Indexes
[0154] The index of visual media constructed by search module 42 may include: attributes of visual media and visual semantic vectors of visual media.
[0155] The search phase includes the following steps:
[0156] S408. Enter the search query.
[0157] Users can access the search interface provided by the gallery service module 41, for example: Figure 3A In interface 301 shown in (a), a search query is entered. For example, as shown in [example text]... Figure 3AThe interface 305 shown in (b) in FIG. 3, and input “surrounding stove tea” in the search box 306.
[0158] S409, sending the search statement.
[0159] After the gallery service module 41 receives the search statement input by the user, the gallery service module 41 sends the search statement to the search module 42 for searching.
[0160] S410, requesting semantic subject identification on the search statement.
[0161] The search module 42 requests the natural language understanding module 44 to perform semantic subject identification on the search statement, to obtain semantic subjects contained in the search statement. The natural language understanding module 44 is based on a natural language understanding model to perform semantic subject identification. Specifically, the search statement can be subjected to semantic subject identification by using a named entity recognition technology (NER), to obtain semantic subjects contained in the search statement. In the embodiments of the present application, the semantic subjects can also be referred to as entities.
[0162] The named entity recognition technology can identify semantic subjects related to time, semantic subjects related to location, and semantic subjects related to tags in the search statement. The semantic subjects related to time and the semantic subjects related to location belong to semantic subjects irrelevant to visual content. The tags of the visual media file are data that can be obtained only by natural picture understanding of a model, and therefore the semantic subjects related to tags belong to semantic subjects relevant to visual content. In actual applications, a plurality of tags that are more concerned by users can be counted according to actual experience, such as “sky”, “cat”, “dog”, “birthday”, “child”, and the like. These tags are used to describe visual content. Subsequently, the named entity recognition technology can match keywords (or search words) in the search statement with the plurality of tags set in advance, to determine whether the keywords belong to semantic subjects related to tags.
[0163] For example, the named entity recognition technology is used to perform semantic subject identification on the search statement “sky shot in Beijing during the National Day”, to determine that “National Day” belongs to a semantic subject related to time, “Beijing” belongs to a semantic subject related to location, and “sky” belongs to a semantic subject related to a tag.
[0164] S411, returning the semantic subjects.
[0165] The natural language understanding module 44 returns the identified semantic subjects to the search module 42.
[0166] S412, requesting semantic understanding on the search statement.
[0167] The search module 42 can send the search sentence to the multi-modal understanding module 43, and the multi-modal understanding module 43 performs semantic understanding on the search sentence to obtain a sentence semantic vector (i.e., a first sentence semantic vector) of the search sentence. The specific semantic understanding process can refer to the corresponding content in the above embodiments, which will not be described here.
[0168] It should be additionally noted that when the search sentence includes a semantic subject related to visual content (i.e., a semantic subject related to a label), the search module 42 can also send the semantic subject related to visual content to the multi-modal understanding module 43 for semantic understanding of the semantic subject related to visual content by the multi-modal understanding module 43 to obtain a subject semantic vector of the semantic subject. Continuing with the above example, “sky” belongs to a semantic subject related to visual content, and the multi-modal understanding module 43 can perform semantic understanding on “sky” to obtain a subject semantic vector corresponding to “sky”.
[0169] S413, return the text vector.
[0170] The multi-modal understanding module 43 can return the sentence semantic vector of the search sentence to the search module 42.
[0171] S414, recall based on the attribute of the visual media and the visual semantic vector of the visual media, respectively.
[0172] The search module 42 includes different branches of search methods:
[0173] For example, the first branch is to recall based on the visual semantic vector of the visual media. Specifically, the visual semantic vector of each visual media in the plurality of visual media stored via the mobile phone is obtained; the vector similarity between the sentence semantic vector of the search sentence and the visual semantic vector of each visual media is calculated; according to the vector similarity, M candidate visual media (i.e., M candidate visual media files) whose visual semantic vector matches the sentence semantic vector are determined from the plurality of visual media stored via the mobile phone; wherein M is an integer greater than or equal to 1; and the M visual media can be taken as a set of visual media recalled by the first branch.
[0174] For example, the visual media with a vector similarity greater than a preset similarity threshold is taken as the visual media whose visual semantic vector matches the sentence semantic vector.
[0175] For example, the plurality of visual media is sorted in descending order of the vector similarity, and the top F (F≥1) visual media is taken as the visual media whose visual semantic vector matches the sentence semantic vector.
[0176] Exemplarily, Q (Q≥1) visual media with vector similarity greater than a preset similarity threshold are determined from the plurality of visual media stored via the mobile phone; if Q is greater than or equal to a preset quantity threshold D, the Q visual media can be sorted in descending order of vector similarity; and the D visual media with the highest ranking are taken as the visual media matched with the semantic vector of the sentence; if Q is less than the preset quantity threshold D, the Q visual media can be directly taken as the visual media matched with the semantic vector of the sentence.
[0177] It should be noted that the plurality of visual media stored via the mobile phone can include visual media stored locally by the mobile phone and / or visual media stored in the cloud (for example, cloud storage space applied for by the mobile phone) by the mobile phone. In order to protect user privacy, the plurality of visual media stored via the mobile phone are all stored in the mobile phone.
[0178] It should be noted that the first branch does not understand the attributes of the visual media in the recall process, and can only understand the overall visual semantic information of the visual media.
[0179] Exemplarily, the second branch is to perform recall based on the attributes of the visual media.
[0180] Specifically, the attributes of each visual media in the plurality of visual media stored via the mobile phone are obtained; the semantic subjects in the search sentence are matched with the attributes of each visual media to determine the visual media (also referred to as second visual media) matched with the semantic subjects; and the visual media matched with the semantic subjects are taken as the visual media set recalled by the second branch. When the number of semantic subjects in the search sentence is one, the visual media set recalled by the second branch includes the visual media matched with the one semantic subject; and when the number of semantic subjects in the search sentence is multiple, the visual media set recalled by the second branch includes the visual media matched with each of the multiple semantic subjects.
[0181] Exemplarily, for a time-related semantic subject, a time range (also referred to as time search range) corresponding to the semantic subject can be determined; the time range is matched with the attribute of the collection time of each visual media to determine the visual media with the collection time located in the time range; and the visual media with the collection time located in the time range are taken as the visual media matched with the semantic subject. For example, the semantic subject is “National Day”, the time range corresponding to the semantic subject is “October 1st to October 7th”, the collection time of picture 1 is “October 2nd”, and the collection time of picture 2 is “October 8th”. According to the above matching method, picture 1 is matched with the semantic subject “National Day”, and picture 2 is not matched with the semantic subject “National Day”.
[0182] Exemplarily, for a semantic subject related to a location, a geographical range corresponding to the semantic subject (i.e. a geographical search range) can be determined; the geographical range is matched with the attribute of the collection location of each visual media to determine the visual media whose collection location is within the geographical range; and the visual media whose collection location is within the geographical range is taken as the visual media matched with the semantic subject. For example, the semantic subject is "Beijing", the geographical range corresponding to the semantic subject is the whole city of Beijing, the collection location of picture 3 is "Xicheng District, Beijing", and the collection location of picture 4 is "Nankai District, Tianjin", then according to the above matching manner, picture 3 is matched with the semantic subject "Beijing", and picture 4 is not matched with the semantic subject "Beijing".
[0183] Exemplarily, for a semantic subject related to a label, the subject semantic vector of the semantic subject and the label semantic vector of the label of each visual media can be obtained; the vector similarity between the subject semantic vector of the semantic subject and the label semantic vector of the label of each visual media is calculated; and the visual media matched with the semantic subject is determined according to the vector similarity. For example, the semantic subject is "human baby", and picture 5 has a label "child", through calculation, it is found that the subject semantic vector of "human baby" is similar to the label semantic vector of "child", i.e. picture 5 is matched with the semantic subject "human baby". In actual application, after obtaining the visual media set recalled by the first branch and the visual media set recalled by the second branch, the candidate visual media set can be determined according to the visual media set recalled by the first branch and the visual media set recalled by the second branch. In an optional implementation, the union or intersection of the visual media set recalled by the first branch and the visual media set recalled by the second branch can be taken as the candidate visual media set.
[0184] In actual application, when a user searches pictures on a mobile phone, sometimes the user pays attention to the visual semantic information of the pictures, sometimes the user pays attention to the attribute information such as the collection location and collection time of the pictures, and sometimes the user pays attention to both. Exemplarily, when a user searches "photos taken today", the user pays attention to the collection time of the pictures; when a user searches "sky taken today", the user pays attention to not only the collection time of the pictures, but also the visual semantic of the pictures, i.e. whether the picture content is the sky; when a user searches "photos taken while walking in Beijing", the user pays attention to the collection location of the pictures; and when a user searches "photos of walking taken in Beijing", the user pays attention to not only the collection location attribute of the pictures, but also the visual semantic of the pictures, i.e. whether the picture content is a walking photo.
[0185] Taking the two search statements "photos taken in Beijing this year" and "sky taken in Beijing this year" as examples, referring to the foregoing introduction, the semantic proportion of visual content in the search statement "sky taken in Beijing this year" is greater than that in the search statement "photos taken in Beijing this year". Obviously, for the search statement "photos taken in Beijing this year", it is more appropriate to take the visual media set recalled by the second branch as the candidate visual media set, so that subsequently, the intersection of the visual media matched by "this year" and the visual media matched by "Beijing" in the candidate visual media set can be taken as the final search result. In the final search result, the collection locations of the visual media are all Beijing and the collection times are all this year, which meets the user's search demand. If the visual media set recalled by the first branch is taken as the candidate visual media set, the following situation may occur: there is a picture of a girl holding a camera to take a picture in the visual media set recalled by the first branch, but the picture is taken in Shanghai and the collection time is last year. Since the search statement "photos taken in Beijing this year" contains the word "take", the picture of the girl holding a camera to take a picture contains the action of "taking", so there is a certain similarity between the sentence semantic vector of the search statement "photos taken in Beijing this year" and the visual semantic vector of the picture of the girl holding a camera to take a picture, that is, the picture of the girl holding a camera to take a picture is likely to be recalled. That is, when the semantic proportion of visual content is low, it is not appropriate to take the visual media set recalled by the first branch as the candidate visual media set.
[0186] To solve the above problem, in an optional embodiment, the semantic proportion of visual content in the search statement can be determined; according to the semantic proportion of visual content in the search statement, the first extraction ratio corresponding to the visual media set recalled by the first branch and the second extraction ratio corresponding to the visual media set recalled by the second branch are determined; when the semantic proportion is greater than or equal to a preset proportion threshold, the first extraction ratio is greater than the second extraction ratio. For example, the semantic proportion is S, the first extraction ratio is β*S, and the second extraction ratio is 1-β*S, where the value of β can be set according to actual needs, and the embodiments of the present application do not make specific limitations thereto.
[0187] In another optional implementation, the semantic proportion related to visual content in the search query can be determined. When the semantic proportion related to visual content in the search query is greater than or equal to a preset threshold, the visual media set recalled by the first branch is used as a candidate visual media set. When the semantic proportion related to visual content in the search query is less than the preset threshold, the visual media set recalled by the second branch is used as a candidate visual media set. A semantic proportion related to visual content in the search query greater than or equal to the preset threshold indicates that the search query is a high-semantic search query; a semantic proportion related to visual content in the search query less than the preset threshold indicates that the search query is a low-semantic search query. The specific process will be described in the following description. Figure 5 The embodiments are described in detail below.
[0188] S415, Visual Media Filtering.
[0189] Search module 42 can perform semantic subject filtering, spatiotemporal filtering, and / or interpersonal relationship filtering on candidate visual media sets. Specific filtering methods will be described in the following sections. Figure 5 The embodiments are described in detail below.
[0190] S416, Visual Media Ordering.
[0191] Search module 42 sorts the candidate visual media sets.
[0192] The specific sorting method will be described in detail in the following examples.
[0193] S417, Return to search results.
[0194] The search module 42 sends the sorted search results to the image library service module 41.
[0195] Specifically, the search module 42 can select the top P (P≥1) visual media and their ranking information as the final search results.
[0196] S418. Display search results.
[0197] Image gallery service module 41 can display search results to users, for example: through Figure 3A Interface 305 Figure 3A The search results are displayed in interface 308.
[0198] like Figure 3A Interface 305 shows 239 images in the search results, but only displays thumbnails of 8 of them. To view the full 239 images, users can click the "More" option on interface 305. In response, the phone displays... Figure 3Athe interface 308 in FIG. 1. In an example, the display order of the pictures in the search results is related to the matching degree between the pictures and the search statement, for example, the matching degree between the pictures displayed in the front and the search statement is greater than or equal to the matching degree between the pictures displayed in the back and the search statement. The calculation of the matching degree and the sorting manner will be described in detail in the following embodiments.
[0199] The following will be described in detail with reference to the accompanying drawings. Figure 5 The search process performed by the search module 42 of the present application will be described in detail:
[0200] 501. Receive a search statement.
[0201] 502. Perform a search process corresponding to the search statement based on the visual semantic vector of the visual media.
[0202] The first branch recalled visual media set is obtained by performing step 502, which can be specifically referred to the search process corresponding to the first branch described above.
[0203] For example, it is assumed that the multi-modal model can encode the visual media and the search statement into a k-dimensional vector space. The search statement is Q, and the corresponding sentence semantic vector V Q = {a z}, z = 1, 2, …, Z, and M pictures are recalled, and the vector representation is:
[0204]
[0205] Wherein, the i-th row corresponds to the visual semantic vector of the i-th picture, and i takes the value range of 【1, M】.
[0206] 503. Semantic subject identification.
[0207] The specific process of semantic subject identification of the search statement can be referred to the corresponding content in the above embodiments, which will not be described here.
[0208] 504. Search based on the attribute of the visual media.
[0209] Specifically, the semantic subject irrelevant to the time content in the search statement is matched with the attribute of the visual media to obtain the second branch recalled visual media set, which can be specifically referred to the search process corresponding to the second branch described above.
[0210] 505. Rewrite the search statement.
[0211] Specifically, the rewritten search statement is obtained by deleting the semantic subject irrelevant to the visual content in the search statement. The semantic subject irrelevant to the visual content specifically refers to the semantic subject related to time and the semantic subject related to place.
[0212] An example is the search sentence "sky shot this year", where "this year" is a time-related semantic subject, and the rewritten search sentence is "sky shot".
[0213] In practical applications, after deleting the semantic subjects irrelevant to visual content, there can be some redundant stop words, for example, the search sentence "sky shot in Beijing this year", where "this year" is a time-related semantic subject and "Beijing" is a location-related semantic subject. After deleting "this year" and "Beijing", the stop word "in" becomes redundant, and thus needs to be deleted. Specifically, the semantic subjects irrelevant to visual content and their related stop words in the search sentence are deleted to obtain a rewritten search sentence. An example is that the search sentence "sky shot in Beijing this year" corresponds to the rewritten search sentence "sky shot".
[0214] It should be noted that the above steps 502, 503 and 505 have no order restrictions in execution. In an optional example, in order to improve efficiency, the three steps can be executed simultaneously.
[0215] 506, performing a search process corresponding to the rewritten search sentence based on the visual semantic vector of the visual media.
[0216] The execution of step 506 obtains a third branch recalled visual media set.
[0217] Specifically, the visual semantic vector of each visual media in the plurality of visual media stored via the mobile phone is obtained; the vector similarity between the sentence semantic vector of the rewritten search sentence (i.e., the second sentence semantic vector) and the visual semantic vector of each visual media is calculated; according to the vector similarity, a plurality of visual media (i.e., reference visual media) whose visual semantic vector matches the sentence semantic vector is determined from the plurality of visual media stored via the mobile phone; and the plurality of visual media can be taken as the third branch recalled visual media set.
[0218] The specific implementation process of the step "determining a plurality of visual media whose visual semantic vector matches the sentence semantic vector from the plurality of visual media stored via the mobile phone according to the vector similarity" can refer to the corresponding content in the above embodiments, which will not be described here.
[0219] An example is that the rewritten search sentence is Q', and the corresponding vector V Q′ z}, z = 1, 2,..., Z, and M' pictures set R' are recalled, and the vector representation is:
[0220]
[0221] where the i-th row corresponds to the visual semantic vector of the i-th picture, and i takes a value in the range [1, M'].
[0222] 507. Calculate the semantic proportion related to the visual content in the search statement.
[0223] The following will introduce a determination manner of the semantic proportion related to the visual content:
[0224] 5071. Determine the representative visual semantic vector V I (that is, the first representative visual semantic vector) corresponding to the visual media set recalled by the first branch and the representative visual semantic vector V I′ (that is, the second representative visual semantic vector) corresponding to the visual media set recalled by the third branch.
[0225] 5072. Determine the difference between the sentence semantic vector V Q and the representative visual semantic vector V I to obtain a difference vector (V Q -V I (that is, the first difference vector).
[0226] 5073. Determine the difference between the sentence semantic vector V Q′ and the representative visual semantic vector V I′ to obtain a difference vector (V Q′ -V I′ (that is, the second difference vector).
[0227] 5074. According to the vector similarity between the difference vector (V Q -V I ) and the difference vector (V Q′ -V I′ ), determine the semantic proportion related to the visual content in the search statement.
[0228] Wherein, the semantic proportion related to the visual content is positively correlated with the vector similarity.
[0229] In the above 5071, in an example, the vector similarity between the visual semantic vector of each visual media in the visual media set recalled by the first branch and the sentence semantic vector of the search statement can be obtained; the visual media in the visual media set recalled by the first branch is sorted according to the vector similarity from high to low; and the average vector of the visual semantic vectors of the top T (T≥1) visual media is taken as the representative visual semantic vector V I .
[0230] Continuing with the above example, the representative visual semantic vector V I is:
[0231]
[0232] The vector similarity between each visual media visual semantic vector in the third branch recalled visual media set and the rewritten search sentence sentence semantic vector can be obtained; the visual media in the third branch recalled visual media set can be sorted according to the vector similarity from high to low; and the average vector of the visual semantic vectors of the top H (H≥1) visual media is taken as the representative visual semantic vector V I′ .
[0233] Continuing with the above example: the representative visual semantic vector V I′ is:
[0234]
[0235] The values of H and T can be the same or different, and embodiments of the present application do not make specific limitations thereon. In another example, the visual media in the first branch recalled visual media set can be clustered using a clustering algorithm according to the vector similarity between the visual semantic vector of each visual media in the first branch recalled visual media set and the search sentence sentence semantic vector, and a clustering center point is obtained; and the visual semantic vector of the clustering center point is taken as the first representative visual semantic vector.
[0236] The visual media in the third branch recalled visual media set can be clustered using a clustering algorithm according to the vector similarity between the visual semantic vector of each visual media in the third branch recalled visual media set and the second search sentence sentence semantic vector of the rewritten search sentence, and a clustering center point is obtained; and the visual semantic vector of the clustering center point is taken as the second representative visual semantic vector.
[0237] In embodiments of the present application, the representative visual semantic vector is used to represent the visual semantic of the entire set. The specific clustering algorithm can be selected according to actual needs, and embodiments of the present application do not make any limitations thereon.
[0238] In an optional embodiment of the above 5074, the vector similarity between the difference vector (V Q -V I ) and the difference vector (V Q′ -VI′) can be directly taken as the semantic proportion related to the visual content in the search sentence.
[0239] Wherein, the greater the vector similarity between the difference vector (V Q -V I ) and the difference vector (V Q′ -V I′ ), the greater the semantic proportion related to the visual content in the search sentence; otherwise, the smaller the semantic proportion related to the visual content in the search sentence.
[0240] It should be noted that in the embodiments of the present application, the calculation method of the semantic proportion of the visual content related semantics in the search statement is based on the part-of-speech characteristics of the distribution representation vector, i.e. the additivity of the word meaning is directly reflected as the additivity of the distribution representation vector.
[0241] When the semantic proportion of the visual content related semantics in the search statement is less than the preset proportion threshold, subsequent step 508 is performed.
[0242] When the semantic proportion of the visual content related semantics in the search statement is greater than or equal to the preset proportion threshold, subsequent steps 509 and 510 are performed.
[0243] 508, the second branch recalled visual media is determined as the candidate set.
[0244] 509, the first branch recalled visual media is determined as the candidate set.
[0245] 510, visual media filtering.
[0246] In order to improve the search accuracy, one or more of the semantic subject filtering, the space-time filtering, the person name filtering and the character relationship filtering can also be performed on the first branch recalled visual media. The semantic subject filtering refers to filtering the first branch recalled visual media by using the matching degree of the first branch recalled visual media and the semantic subject related to the visual content in the search statement. The space-time filtering includes the time filtering and the space (i.e. location) filtering. The person name filtering refers to filtering the first branch recalled visual media by using the matching degree of the person name attribute of the first branch recalled visual media and the semantic subject related to the person name in the search statement. The character relationship filtering refers to filtering the first branch recalled visual media by using the matching degree of the character relationship attribute of the first branch recalled visual media and the semantic subject related to the character relationship in the search statement. In actual application, in addition to the above-mentioned collection location, collection time, visual media name and other attributes, the visual media stored in the mobile phone can also include the person name attribute and the character relationship attribute manually input by the user. Therefore, in actual application, the named entity recognition technology can also be used to match the keywords in the search statement with the preset multiple character relationships to determine whether the keywords belong to the semantic subject related to the character relationship.
[0247] When multiple processing of the semantic subject filtering, the space-time filtering, the person name filtering and the character relationship filtering need to be performed on the first branch recalled candidate set, the execution order of the multiple processing can be set according to actual needs, and the embodiments of the present application do not make specific limitation on this.
[0248] For example, as shown in FIG. 6, the visual media filtering includes: Figure 5 the semantic subject filtering, the space-time filtering, the person name filtering and the character relationship filtering.
[0249] 5101. Semantic subject filtering.
[0250] 5102. Spatio-temporal filtering.
[0251] 5103. Person relationship filtering.
[0252] 5104. Name filtering.
[0253] In the above 5101, generally, when a user searches, he / she wants to recall pictures containing some visual content described in the search sentence, such as sky, puppy, child, etc. That is, the semantic subject related to the visual content in the search sentence represents the visual focus of the user when searching.
[0254] However, the first branch considers the matching degree between the visual content of the picture / video and the entire search sentence when recalling, without fully considering the role of certain specific semantic subjects (i.e., semantic subjects related to visual content) in the search sentence in the mobile phone gallery search scenario, resulting in the recall of some inaccurate photos.
[0255] For example, when a user searches for "photos of the Great Wall taken during the National Day", the semantic subject related to the visual content includes "the Great Wall". When the above first branch makes a match, it considers the matching degree between the entire search sentence and the visual content of the picture, which may result in the first branch recalling a picture containing the "taking" behavior but not the "Great Wall". This is because the picture contains the "taking" behavior, the search sentence contains the word "taking", that is, there is a certain similarity between the visual semantic vector of the picture and the sentence semantic vector of the search sentence, so the picture has the possibility of being recalled.
[0256] For another example, when a user searches for "last year's child holding a cake on his / her birthday", the semantic subject related to the visual content includes "child", "birthday", and "cake". When the first branch makes a match, it considers the matching degree between the entire search sentence and the picture, which may result in the model recalling a picture of an adult holding a cake on his / her birthday. This is because the visual semantic vector of the picture of the adult holding a cake on his / her birthday is highly similar to the sentence semantic vector of "last year's child holding a cake on his / her birthday".
[0257] Therefore, in order to further improve the accuracy of searching visual media in a mobile phone, the "semantic subject" can be strengthened on the basis of the first branch recalling the visual media, that is, the matching degree between the visual media and the "semantic subject" is used to fine-tune the results of the first branch.
[0258] Specifically, the "semantic subject filtering" in the above 5101 can include the following steps:
[0259] 5101a, determine the dimensions to be matched according to the semantic subjects related to the visual content in the search statement.
[0260] 5101b, filter the M candidate visual media according to the matching degrees of the visual content of the M candidate visual media in the dimensions to be matched.
[0261] In an example, the semantic subjects related to the visual content are taken as the dimensions to be matched in 5101a. When there are multiple semantic subjects related to the visual content in the search statement, the multiple semantic subjects are taken as different dimensions to be matched respectively to obtain multiple dimensions to be matched. The number of the multiple dimensions to be matched is consistent with the number of the multiple semantic subjects related to the visual content.
[0262] For example, the search statement "last year, the child held the cake for birthday" has the semantic subjects "child", "birthday" and "cake" related to the visual content, and the corresponding three dimensions to be matched are "child", "birthday" and "cake" respectively.
[0263] In practical applications, besides taking the semantic subjects as the dimensions to be matched, the search statement itself can also be taken as the dimension to be matched, so that the matching of the visual content of the candidate visual media with the whole search statement can be considered when filtering the semantic subjects, to improve the rationality of the filtering. Specifically, the semantic subjects related to the visual content and the search statement can be taken as different dimensions to be matched respectively to obtain multiple dimensions to be matched. When there are multiple semantic subjects related to the visual content in the search statement, the multiple semantic subjects and the search statement are taken as different dimensions to be matched respectively. The number of the multiple dimensions to be matched is one more than the number of the multiple semantic subjects related to the visual content.
[0264] For example, the search statement "last year, the child held the cake for birthday" has the semantic subjects "child", "birthday" and "cake" related to the visual content, and the corresponding four dimensions to be matched are "child", "birthday", "cake" and "last year, the child held the cake for birthday" respectively.
[0265] In the 5101b, the matching degree of the visual content of the candidate visual media on the to-be-matched dimension refers to the matching degree of the visual content of the candidate visual media and the to-be-matched dimension. The vector similarity between the visual semantic vector of the candidate visual media and the to-be-matched semantic vector of the to-be-matched dimension can be determined; when the to-be-matched dimension is a semantic subject related to the visual content, the to-be-matched semantic vector of the to-be-matched dimension is the subject semantic vector of the semantic subject; when the to-be-matched dimension is a search statement, the to-be-matched semantic vector of the to-be-matched dimension is the sentence semantic vector of the search statement; according to the vector similarity, the matching degree of the visual content of the candidate visual media on the to-be-matched dimension is determined. The matching degree is positively correlated with the vector similarity.
[0266] When the number of to-be-matched dimensions is one, the candidate visual media with a matching degree less than or equal to a preset matching degree threshold can be filtered out according to the matching degrees of the visual content of the plurality of candidate visual media on the to-be-matched dimension.
[0267] For example, the search statement is "take a picture of the sky on National Day", and the only semantic subject related to the visual content is "sky"; then, the recall picture containing the "take" behavior but not containing the "sky" has a relatively low matching degree with the "sky" and will be filtered out.
[0268] When the number of to-be-matched dimensions is multiple, for each candidate visual media, the comprehensive matching degree of the candidate visual media can be determined according to the matching degrees of the visual content of the candidate visual media on the multiple to-be-matched dimensions; and the plurality of candidate visual media can be filtered according to the comprehensive matching degrees of each candidate visual media. Specifically, from the M candidate visual media, a plurality of first visual media with a comprehensive matching degree greater than or equal to a preset matching degree threshold (i.e., meeting a preset requirement) are determined, which is equivalent to filtering out the candidate visual media with a comprehensive matching degree less than the preset matching degree threshold; or, the M candidate visual media are sorted in descending order of the comprehensive matching degrees, and the Z' (Z'≥1) candidate visual media at the front of the sorting (i.e., meeting the preset requirement) are taken as the plurality of first visual media, which is equivalent to filtering out the (M-Z') candidate visual media at the rear of the sorting.
[0269] In an optional implementation, the comprehensive matching degree of the candidate visual media can be determined in any one of the following three ways:
[0270] The first way is to sum the matching degrees of the visual content of the candidate visual media on the multiple to-be-matched dimensions to obtain the comprehensive matching degree of the candidate visual media.
[0271] The second way is to perform weighted summation on the matching degrees of the visual content of the candidate visual media on the multiple to-be-matched dimensions to obtain the comprehensive matching degree of the candidate visual media.
[0272] The weight of each of the plurality of to-be-matched dimensions can be configured by a user in advance.
[0273] In a third mode, a machine learning model is used to determine the comprehensive matching degree of the candidate visual media according to the matching degrees of the visual content of the candidate visual media in the plurality of to-be-matched dimensions.
[0274] The machine learning model needs to be trained based on a data set, and the purpose of the training is also essentially to learn the weight of each to-be-matched dimension.
[0275] In the first mode, the contribution of the matching degrees in different dimensions to the comprehensive matching degree is not distinguished, which may result in that the numerical values of the comprehensive matching degrees of the plurality of candidate visual media are close or equal, and thus the plurality of candidate visual media cannot be screened.
[0276] For example, it is assumed that the plurality of candidate visual media includes picture 1, picture 2 and picture 3, and the plurality of to-be-matched dimensions includes dimension A, dimension B and dimension C. The matching degrees of the visual content of each picture in each dimension are calculated respectively, and the results are shown in Table 1.
[0277] Table 1:
[0278] Dimension A Dimension B Dimension C Picture 1 0.30 0.35 0.55 Picture 2 0.40 0.38 0.42 Picture 3 0.38 0.37 0.45
[0279] If the first mode is used for calculation, the comprehensive matching degrees of picture 1, picture 2 and picture 3 are all 1.2, which results in that the three pictures cannot be screened.
[0280] In the second mode, when the number of the semantic subjects related to the visual content that need to be focused on is too large (i.e., the number of the preset plurality of labels is too large), it is difficult to accurately configure the weights of different dimensions.
[0281] In the third mode, the construction of the data set cannot be implemented without user data. However, the user data belongs to user privacy content, and the user does not want to report the data to the cloud side.
[0282] In an optional embodiment, in order to solve the above problems, the following steps can be used to determine the comprehensive matching degree:
[0283] S51, determining the weight of each of the N to-be-matched dimensions.
[0284] The weight of the jth to-be-matched dimension is positively correlated with the variation degree of the matching degree of the visual content of the M candidate visual media files in the jth to-be-matched dimension; j is an integer, and the value of j is sequentially from 1 to N;
[0285] S52, according to the weight of the N dimensions to be matched, the matching degrees of the visual content of the i-th candidate visual media file on each dimension to be matched are weighted and summed to obtain a comprehensive matching degree of the i-th candidate visual media file.
[0286] wherein i is an integer, and i takes values from 1 to M in turn.
[0287] Taking the search statement "last year, the child held a cake for birthday" as an example, the semantic subjects related to the visual content include: child, birthday, and cake. If multiple recalled photos all contain a cake, then the variation degree of the matching degrees of the visual content of the multiple recalled photos on the dimension of "cake" will be relatively small, and the weight corresponding to the dimension of "cake" will be relatively small. If multiple recalled photos contain a child or do not contain a child, then the variation degree of the matching degrees of the visual content of the multiple recalled photos on the dimension of "child" will be relatively large, and the weight corresponding to the dimension of "child" will be relatively large.
[0288] In an example, for each dimension to be matched, the variation degree of the matching degrees of the visual content of the multiple candidate visual media on the dimension to be matched is determined according to the information entropy of the matching degrees of the visual content of the multiple candidate visual media on the dimension to be matched; wherein the variation degree is inversely proportional to the information entropy. The calculation method of the information entropy will be described in detail in the following embodiments.
[0289] In this embodiment, the information entropy is used to measure the variation degree of the matching degrees on each dimension to be matched, and the weight corresponding to the dimension to be matched is determined based on the variation degree.
[0290] In order to ensure that the matching degrees of the visual content of the candidate visual media on different dimensions to be matched have a unified dimension, a normalization processing step can be performed. Specifically, for each candidate visual media, the initial matching degree of the candidate visual media on the dimension to be matched is determined according to the vector similarity between the visual semantic vector of the candidate visual media and the matching semantic vector of the dimension to be matched. For example, the vector similarity between the visual semantic vector of the candidate visual media and the matching semantic vector of the dimension to be matched can be taken as the initial matching degree of the candidate visual media on the dimension to be matched. The initial matching degrees of the visual content of the multiple candidate visual media on the dimension to be matched are normalized to obtain the matching degrees of the visual content of the multiple candidate visual media on the dimension to be matched.
[0291] The normalization process and the calculation process of the weight and the comprehensive matching degree will be described in detail below:
[0292] Suppose there are m candidate visual media, n to-be-matched dimensions, and the initial matching degrees of the visual content of the m candidate visual media on each to-be-matched dimension are regarded as a data matrix:
[0293] X = (x ij ) m*n , i = 1, 2, …, m; j = 1, 2, …, n (5)
[0294] wherein x ij is the initial matching degree of the visual content of the i th candidate visual media on the j th to-be-matched dimension.
[0295] Continuing with the above example, the multiple candidate visual media are picture 1, picture 2 and picture 3, and the multiple to-be-matched dimensions are dimension A, dimension B and dimension C, that is, m is 3 and n is 3.
[0296] First step: normalizing the above data matrix.
[0297] The normalized matrix is:
[0298] R = (r ij ) m*n , i = 1, 2, …, m; j = 1, 2, …, n (6)
[0299] wherein,
[0300]
[0301] wherein max(x j ) is the maximum value of the initial matching degrees of the visual content of the multiple candidate visual media on the j th to-be-matched dimension; and min(x j ) is the minimum value of the initial matching degrees of the visual content of the multiple candidate visual media on the j th to-be-matched dimension.
[0302] In actual applications, other normalization methods can also be used, which are not specifically limited in the embodiments of the present application.
[0303] For example, the results obtained by normalizing the example data in Table 1 are shown in Table 2.
[0304] Table 2:
[0305] Dimension A Dimension B Dimension C Picture 1 0.00 0.00 1.00 Picture 2 1.00 1.00 0.00 Picture 3 0.80 0.67 0.23
[0306] Second step: calculating the information entropy corresponding to each to-be-matched dimension.
[0307] The formula (8) and formula (9) can be used for calculation:
[0308]
[0309] wherein,
[0310]
[0311] wherein, e j denotes the information entropy corresponding to the jth matching dimension.
[0312] wherein, since the domain of the function ln(x) is x>0. In actual calculation, in order to avoid the case that p ij in ln(p ij ) is 0, ln(p ij ) in formula (8) can be replaced by ln(p ij +α), wherein α<<0.001.
[0313] The smaller the information entropy corresponding to the matching dimension is, the greater the variation degree of the matching degree of the visual content of the multiple candidate visual media in the matching dimension is, and the greater the information quantity provided is. It can be considered that the matching dimension plays a greater role in the comprehensive evaluation.
[0314] For example, according to the information entropy calculated by formula (8) and formula (9) for the example data in table 2, the information entropy is shown in table 3:
[0315] Table 3:
[0316] Dimension A Dimension B Dimension C Information Entropy 0.69 0.67 0.48
[0317] Step 3: Calculate the weight corresponding to each matching dimension.
[0318] Formula (10) can be used to calculate the weight corresponding to each matching dimension:
[0319]
[0320] wherein, d j denotes the weight corresponding to the jth matching dimension.
[0321] It can be seen that formula (10) is a monotone decreasing function of the information entropy. In this embodiment, the smaller the information entropy is, the greater the variation degree is; the greater the variation degree is, the greater the weight is. That is, the weight is negatively correlated with the information entropy.
[0322] In order to ensure that the sum of the weights corresponding to the multiple matching dimensions is 1, formula (11) can be used to calculate the final weight w j as follows:
[0323]
[0324] wherein, w jdenotes the maximum weight corresponding to the jth dimension to be matched.
[0325] In an optional embodiment, other monotonically decreasing functions can also be used to calculate the above-mentioned weights, which are not limited in the embodiments of the present application.
[0326] For example, for the example data in Table 3, the weights of different dimensions to be matched can be calculated according to the above-mentioned formulas (10) and (11), as shown in Table 4:
[0327] Table 4:
[0328] Dimension A Dimension B Dimension C Weight 0.28 0.29 0.42
[0329] It should be noted that, because the rounding processing is introduced in the process of calculating the information entropy, the sum of the three dimensions in Table 4 above is not 1.
[0330] Step 4: Weighted summation.
[0331] The normalized matching degrees of any candidate visual media are weighted and summed to obtain the comprehensive matching degree of the candidate visual media, which can be calculated according to formula (12):
[0332]
[0333] Wherein, s i denotes the comprehensive matching degree of the i th visual media.
[0334] For example, for the example data in Table 2 and Table 4, the comprehensive matching degrees of different pictures can be calculated according to the above-mentioned formula (12), as shown in Table 5:
[0335] Table 5:
[0336] Picture 1 Picture 2 Picture 3 Comprehensive Matching Degree 0.42 0.58 0.52
[0337] In 5102 above, the first branch only considers the visual semantic information of the visual media when recalling the visual media, without considering the attribute information such as time and location of the visual media. Therefore, the visual media recalled by the first branch needs to be filtered by time and location. Specifically, when the user performs a semantic search, if the search statement contains time information, the pictures that do not meet the time limit need to be filtered out.
[0338] However, the user's description of time information is diverse and fuzzy. For example, the user searches for "sky shot in the afternoon", and the "afternoon" in the user's search statement does not have a clear and standard definition. When the user uses a fuzzy time expression in the search statement, the time window used by the time filtering will affect the user's experience.
[0339] For example, the user starts a trip to a place outside the home city on September 30, 2022, and returns home on October 9, 2022. During this period, the user takes a lot of photos. On a certain day in 2023, the user wants to view the beautiful scenery taken during the trip. The user is likely to input the search statement “scenery taken during the National Day holiday last year”. If the time window for time filtering is set to “from October 1, 2022 to October 7, 2022” according to “last year National Day” in the search statement, the scenery photos taken by the user on September 30, 2022, October 8, 2022, and October 9, 2022 will be filtered out, which obviously does not meet the user's expectation. To improve the rationality of time filtering, the embodiments of the present application provide a new time filtering method. Specifically, the clustering algorithm is used to cluster the plurality of candidate visual media according to the collection time of each of the plurality of candidate visual media, to obtain K (K≥1) clustering clusters.
[0340] In this way, the same series of pictures with close collection times can be classified into the same clustering cluster.
[0341] Continuing with the above example, the natural scenery picture taken by the user on September 30, 2022 and the natural scenery picture taken by the user on October 1, 2022 have close collection times, and are classified into the same clustering cluster by the above clustering algorithm.
[0342] The above clustering algorithm can include but is not limited to K-Means clustering algorithm, Mean shift clustering algorithm, and density-based clustering algorithm.
[0343] Taking the density-based clustering algorithm as an example, in the scenario of semantic search containing time information, it is unreasonable to set a fixed first parameter ∈ and second parameter MinPts. For example, in the case of the user searching for “photos taken during the trip in 2022”, the time range is one whole year; and in the case of the user searching for “photos taken during the morning trip”, the time range is a few hours. The same first parameter ∈ and second parameter MinPts should not be set in these two cases. To improve the rationality of clustering, the following steps can be used to determine the first parameter ∈ and the second parameter MinPts:
[0344] 51021、According to the time search range contained in the search statement, determine the first parameter involved in the density-based clustering algorithm.
[0345] The first parameter is positively correlated with the time length corresponding to the time search range.
[0346] The time search range can be determined according to the time-related semantic subject in the search statement. Specifically, the time range corresponding to the time-related semantic subject (i.e., the time search range) can be returned through a mapping table. The mapping table can be constructed in advance as needed, and the specific form is not limited in the embodiments of the present application.
[0347] For example, the semantic subject related to time extracted from the search statement "photos taken in spring" is "spring", the mapping table is queried, and the corresponding time range is obtained as: February 1 to May 30; the semantic subject related to time extracted from the search statement "sky taken in the morning" is "morning", the mapping table is queried, and the corresponding time range is obtained as: 7:00-12:00; the semantic subject related to time extracted from the search statement "photos taken for the National Day" is "National Day", the mapping table is queried, and the corresponding time range is obtained as: October 1 to October 7; the semantic subject related to time extracted from the search statement "sky taken in Beijing this year" is "this year", the mapping table is queried, and the corresponding time range is obtained as "from January 1, 2023 to December 31, 2023".
[0348] The first parameter can be determined according to the duration corresponding to the time search range. For example, the start time (which can be understood as the start time stamp) of the time search range is T start , and the end time (which can be understood as the end time stamp) of the time search range is T end , and the duration of the time search range is: T end -T start The first parameter can be calculated by the following formula:
[0349] ∈=α*(T end -T start ) (13)
[0350] Wherein, α is a coefficient that can be artificially adjusted, and its size can be set according to actual needs, and the present application does not make specific limitation.
[0351] 51022, the second parameter involved in the density-based clustering algorithm is determined according to the ratio of the number of multiple candidate visual media to the time search range.
[0352] The second parameter can be calculated by the following formula:
[0353]
[0354] N is the total number of visual media in the candidate set whose collection time is within the time search range, N≥1; n is the number of times the time search range repeats between T1 and T2. T1 is the collection time of the earliest collected visual media in the plurality of visual media stored via the mobile phone; T2 is the collection time of the latest collected visual media in the plurality of visual media stored via the mobile phone. For example, the search statement is "sky shot during the National Day", the time search range is from October 1 to October 7, the time length is 7 days, the collection time of the earliest collected visual media stored in the mobile phone is August 1, 2020, and the collection time of the latest collected visual media stored in the mobile phone is October 20, 2023. Then, from August 1, 2020 to October 20, 2023, the time search range from October 1 to October 7 repeats 4 times (i.e., once a year).
[0355] (T end -T start It can be understood that the total time length corresponding to the time search range.
[0356] Wherein, β is a coefficient that can be adjusted artificially, and its size can be set according to actual needs, which is not limited in the present application.
[0357] In this embodiment, the first parameter ∈ and the second parameter MinPts involved in the density-based clustering algorithm are dynamically adjusted according to the time length defined by the time search range contained in the search statement, so that the rationality of clustering is ensured, and the accuracy of the final search result is improved.
[0358] When the search statement contains a time search range, determine the collection time range corresponding to each of the K (K≥1) clustering clusters; filter out the clustering clusters whose collection time range does not overlap with the time search range, and retain the clustering clusters whose collection time range overlaps with the time search range.
[0359] The overlap between the collection time range and the time search range can be partial overlap or total overlap, and both belong to the overlap part.
[0360] In this way, G (G≥1) clustering clusters are selected from the K clustering clusters. The collection time of the earliest collected visual media in the G clustering clusters is before the start time of the time search range, and / or the collection time of the latest collected visual media in the G clustering clusters is after the end time of the time search range. Note that there is no intersection between the G clustering clusters, and there is also no intersection between the collection time ranges of the G clustering clusters.
[0361] The collection time range corresponding to the clustering cluster is the collection time T of the earliest collected visual media in the clustering cluster.min T1 to the collection time T of the latest collected visual media in the cluster max , that is: min , T max ].
[0362] Specifically, the time search range is [T start , T end ], the collection time range of the cluster is [T min , T max ], when [T min , T max ] and [T start , T end ] have an overlapping part, the cluster is retained; when [T min , T max ] and [T start , T end ] have no overlapping part, the cluster is filtered. For example, the time search range is October 1, 2022 to October 7, 2022, and the collection time range of the cluster is September 30, 2022 to October 1, 2022, and the two ranges have an overlapping part (i.e. October 1, 2022), and the cluster is retained.
[0363] In the above example, the natural scenery photo taken by the user on September 30, 2022 and the natural scenery photo taken by the user on October 1, 2022 are close in shooting time, and are classified into the same cluster by the above clustering algorithm. Assuming that this cluster contains only these two pictures, the collection time range of the cluster is September 30, 2022 to October 1, 2022, and there is an overlapping part between the time search range October 1, 2022 to October 7, 2022, so the cluster is retained, that is, the natural scenery photo taken by the user on September 30, 2022 will not be filtered out.
[0364] Similarly, the natural scenery photo taken by the user on October 8, 2022 and the natural scenery photo taken by the user on October 7, 2022 are close in shooting time, and are classified into the same cluster by the above clustering algorithm. Assuming that this cluster contains only these two pictures, the collection time range of the cluster is October 7, 2022 to October 8, 2022, and there is an overlapping part between the time search range October 1, 2022 to October 7, 2022, so the cluster is retained. That is, the natural scenery photo taken by the user on October 8, 2022 will not be filtered out.
[0365] In this way, for the search statement "scenery shot during the National Day holiday last year", the collection time of the earliest collected visual media in the multiple clustering clusters filtered is September 30, 2022, and the collection time of the latest collected visual media is October 8, 2022.
[0366] It can be seen that the time filtering method provided by the embodiments of the present application can ensure that the same series of photos with relatively close collection times are displayed to the user, guarantee the continuity of the search results, and improve the search experience of the user.
[0367] In addition, when the search statement also contains a semantic subject related to a location, the G clustering clusters filtered can also be further filtered by location. Specifically, the visual media in the clustering cluster whose collection location does not match the semantic subject related to the location in the search statement can be filtered out. For example, the collection location "Beijing Xicheng Prince Gong's Mansion" can be considered to match the semantic subject "Beijing Haidian Sujiatuo Town" (belonging to city-level matching); the collection location "Beijing Haidian Tsinghua University" can be considered to match the semantic subject "Beijing Haidian Sujiatuo Town" (belonging to district-level matching); and the collection location "Shanghai" can be considered to not match the semantic subject "Beijing".
[0368] In the above embodiments, clustering is performed first, time filtering is then performed, and finally location filtering is performed. Of course, in actual applications, location filtering can also be performed first, clustering can then be performed, and finally time filtering can be performed. The specific execution order of the three steps can be set according to actual needs, and the embodiments of the present application do not make specific limitations in this regard.
[0369] 5103, person relationship filtering.
[0370] When the search statement contains a semantic subject related to a person relationship, the visual media in each clustering cluster whose person relationship attribute does not match the semantic subject is filtered out. For example, the search statement contains "good friend", and the name attribute of the person in picture B is "colleague", which do not match, and picture B is filtered out.
[0371] 5104, name filtering.
[0372] When the search statement contains a semantic subject related to a name, the visual media in each clustering cluster whose name attribute does not match the semantic subject is filtered out. For example, the search statement contains "Zhang San", and the name attribute of the person in picture A is "Li Si", which do not match, and picture A is filtered out.
[0373] It should be noted that the semantic subject related to a name and the semantic subject related to a person relationship in the search statement can also be identified by a named entity recognition technology. The name attribute of the visual media and the person relationship attribute are manually added to the visual media by the user in advance.
[0374] 511. Sort.
[0375] Based on the candidate set obtained in step 508 above, the visual media in the candidate set can be sorted according to the acquisition time to obtain the display order of the visual media. The visual media displayed earlier in the display order were acquired earlier than the visual media displayed later in the display order. Subsequently, the mobile phone can display the candidate set according to the display order of the visual media in the candidate set.
[0376] The filtered candidate set obtained from the above 510 steps can be sorted as follows:
[0377] The filtered candidate set includes the aforementioned G clusters. The G clusters are then sorted according to the start time of their collection time ranges to obtain their display order (i.e., inter-cluster sorting). For example, as shown... Figure 6 As shown in interface 601, cluster A corresponds to a collection time range from September 30, 2022 to October 1, 2022; cluster B corresponds to a collection time range from October 3, 2022 to October 5, 2022. Therefore, the display order of cluster A precedes the display order of cluster B. For each cluster, the visual media within the cluster are sorted from high to low according to their target matching degree, resulting in the display order among the visual media within the cluster (intra-cluster sorting); or, for each cluster, the visual media within the cluster are sorted according to their collection time, resulting in the display order among the visual media within the cluster. Subsequently, the mobile phone displays the visual media from G clusters based on the inter-cluster sorting and intra-cluster sorting.
[0378] Furthermore, in terminal devices, taking images as an example, when users search for images, they not only focus on the visual information but also on the image's attribute information (e.g., the location and time of the photo). Therefore, how to fuse and sort the recalled images is particularly important. Assume that the candidate set obtained in step 508 or the filtered candidate set obtained in step 510 includes Y visual media; where Y is an integer greater than 1. The following steps can be used to sort the Y visual media to obtain the final search results. Figure 7 As shown, it includes:
[0379] 5111. Based on the search query, determine W dimensions to be matched.
[0380] Where W is an integer greater than 1.
[0381] 5112. Obtain the matching degree and matching degree ranking of Y visual media files on the w-th dimension to be matched.
[0382] Y is an integer greater than 1; w is an integer and takes value from 1 to W.
[0383] 5113、according to the matching degree of the yth visual media file in each of the W dimensions to be matched, correct the total number of times the u th appearance of the yth visual media file, to obtain the corrected total number of times the u th appearance of the yth visual media file.
[0384] wherein u is an integer and takes value from 1 to Y; y is an integer and takes value from 1 to Y.
[0385] 5114、according to the corrected total number of times the yth visual media file appears in each rank, determine the score of the yth visual media.
[0386] 5115、according to the scores of the Y visual media files, sort the Y visual media files.
[0387] In the above 5111, the determination manner of the dimensions to be matched includes multiple manners as follows:
[0388] determining the search statement as the dimension to be matched;
[0389] determining the semantic subject irrelevant to visual content in the search statement as the dimension to be matched;
[0390] determining the user click number as the dimension to be matched.
[0391] In an example, the W dimensions to be matched can include: the search statement itself, the semantic subject related to time in the search statement, the semantic subject related to place in the search statement, and the user click number. The user click number reflects the user's personalized preference. The more the user click number of a certain visual media, the more the user prefers the visual media.
[0392] In the above 5112, the W dimensions to be matched include: the user click number; according to the number of times the yth visual media file has been clicked by the user in history, determine the matching degree of the yth visual media file in the user click number.
[0393] The W dimensions to be matched include: the search statement; the matching degree of the yth visual media file in the search statement refers to the matching degree of the visual semantic vector of the yth visual media file and the first sentence semantic vector of the search statement.
[0394] The W dimensions to be matched include: a semantic subject in the search statement that is irrelevant to the visual content; and determining a matching degree of the yth visual media file on the semantic subject in the search statement that is irrelevant to the visual content according to a matching condition of the semantic subject that is irrelevant to the visual content and an attribute of the yth visual media file. The semantic subject that is irrelevant to the visual content includes a time-related semantic subject and / or a location-related semantic subject. Specifically, the matching degree of the yth visual media file on the time-related semantic subject in the search statement is determined according to a matching condition of the time-related semantic subject and the attribute of the yth visual media file. The matching degree of the yth visual media file on the location-related semantic subject in the search statement is determined according to a matching condition of the location-related semantic subject and the attribute of the yth visual media file.
[0395] The ranking of the matching degrees of the Y visual media files on the wth dimension to be matched is obtained by sorting the matching degrees of the Y visual media files on the wth dimension to be matched from large to small.
[0396] In the above 5113, the total number of times that the yth visual media file appears as the u th is corrected according to the matching degrees of the yth visual media file on each dimension to be matched in the W dimensions to be matched. That is, the corrected total number of times not only considers the real total number of times, but also considers the size of the matching degree, which is helpful to optimize the subsequent ranking effect.
[0397] For example, the matching degree of the 1st visual media file on the 1st dimension to be matched is 1.0 (the optimal matching degree is 1.0), and the 1st visual media file is the 1st. The matching degree of the 1st visual media file on the 2nd dimension to be matched is 0.5, and the 1st visual media file is the 1st. Since the matching degree of the 1st visual media file on the 1st dimension to be matched is large, the value or contribution of the 1st visual media file on the 1st dimension to be matched as the 1st should be large. Since the matching degree of the 1st visual media file on the 2nd dimension to be matched is small, the value or contribution of the 1st visual media file on the 2nd dimension to be matched as the 1st should be small. Therefore, the total number of times is corrected according to the size of the matching degree, which can optimize the subsequent ranking effect.
[0398] In an implementable scheme, the above 5113 can include:
[0399] S61, determining a first weight of the wth dimension to be matched for the yth visual media file according to the matching degree of the yth visual media file on the wth dimension to be matched.
[0400] S62, according to the first weight of each of the W dimensions to be matched for the yth visual media file, weighting and summing the number of times that the yth visual media file appears in the u-th place in each dimension to be matched in the W dimensions to be matched, to obtain the modified total number of times that the yth visual media file appears in the u-th place.
[0401] In an example, in the above S61, the matching degree of the yth visual media file in the w-th dimension to be matched can be determined as the first weight of the w-th dimension to be matched for the yth visual media file.
[0402] In another example, a preset weight of the w-th dimension to be matched is obtained; according to the preset weight of the w-th dimension to be matched, the matching degree of the yth visual media file in the w-th dimension to be matched is modified to obtain a modified matching degree of the yth visual media file in the w-th dimension to be matched; and the modified matching degree of the yth visual media file in the w-th dimension to be matched is determined as the first weight of the w-th dimension to be matched for the yth visual media file.
[0403] The preset weights of the W dimensions to be matched can be set based on prior knowledge, which is not limited in the embodiments of the present application.
[0404] The product of the matching degree of the yth visual media file in the w-th dimension to be matched and the preset weight of the w-th dimension to be matched can be determined as the modified matching degree of the yth visual media file in the w-th dimension to be matched.
[0405] It should be noted that the first weight of the w-th dimension to be matched for different visual media files can be different.
[0406] In the above S62, the number of times that the yth visual media file appears in the u-th place in the w-th dimension to be matched is 1 or 0.
[0407] In addition, before performing the above step 5113, the matching degrees of the Y visual media files in the w-th dimension to be matched can also be processed by percentage.
[0408] In the above 5114, according to the modified total number of times that the yth visual media file appears in each place and the place score of each place, the score of the yth visual media file is determined.
[0409] In the above 5115, the Y visual media files are sorted according to the scores of the Y visual media files from high to low to obtain the display order of the Y visual media files. Subsequently, the visual media files are displayed according to the display order.
[0410] The following will exemplify the calculation process of the score of the visual media file:
[0411] Suppose there are m dimensions and n pictures, the following is the matching score of the i-th picture (1≤i≤n) in the j-th dimension (1≤j≤m).
[0412] For example, suppose the user inputs the search statement "sky shot in Sujiatuo Town, Haidian, Beijing this year", and a total of 4 pictures are recalled; and the search statement "sky shot in Sujiatuo Town, Haidian, Beijing this year" corresponds to the four dimensions of search statement, this year, Beijing Haidian Sujiatuo Town, and click count. The matching of each picture in each dimension is counted respectively, and the results are shown in Table 6.
[0413] Table 6:
[0414] Search Statement This Year Beijing Haidian Sujiatuo Town Clicks Picture 1 0.30 Match Information Missing 50 Picture 2 0.40 No Match No Match 0 Picture 3 0.38 No Match Province and City Match 35 Picture 4 0.50 Match Street Match 50
[0415] The data in Table 6 can be numerically valued to obtain the matching degree of each picture in each dimension, as shown in Table 7.
[0416] Table 7:
[0417] Search Statement This Year Beijing Haidian Sujiatuo Town Clicks Picture 1 0.30 1 0.2 50 Picture 2 0.40 0 0 0 Picture 3 0.38 0 0.8 35 Picture 4 0.50 1 1 50
[0418] 1. Rank score conversion
[0419] The "rank" score vector R=(r1, r2, …, rn) is calculated, where rh= n-h+1, 1≤h≤n, and rhis the score of the h-th picture. h T h
[0420] For example, for the case of 4 pictures, the rank can be converted to obtain:
[0421] Table 8:
[0422] 1st Place 2nd Place 3rd Place 4th Place Score 4 3 2 1
[0423] 2. Percentile processing of matching degree
[0424] The percentile processing of the matching degree of the i-th picture in the j-th dimension is calculated
[0425]
[0426] wherein is the matching degree of the i-th picture in the j-th dimension; represents the optimal matching score in the j-th dimension,
[0427] For example, if the optimal matching scores of the search statement, time, and place are all 1.0, and the optimal matching score of the click times is 100. The percentage of the picture information is calculated according to the above formula as follows:
[0428] Table 9:
[0429] Search Statement This Year Beijing Haidian Sujiatuo Town Clicks Picture 1 0.30 1 0.2 0.50 Picture 2 0.40 0 0 0 Picture 3 0.38 0 0.8 0.35 Picture 4 0.50 1 1 0.50
[0430] 3. Preset weight configuration of each dimension
[0431] Based on prior knowledge, etc., the preset weight of each dimension is configured: W = (w (1) , w (2) , …, w (m) ) T .
[0432] For example, based on expert experience, it is considered that the importance of the time matching degree is twice that of other dimensions, and thus W = (1, 2, 1, 1) T .
[0433] 4. Fusion scoring
[0434] The fusion score of the picture i is calculated as follows:
[0435] s i = R T f i = R T δ i U i W (16)
[0436] Wherein, R T represents the transposition operation on R; represents the ranking matrix of the picture in different dimensions. Specifically, if the i-th picture ranks the h-th in the j-th dimension, then
[0437] For example, for the picture 1, the respective ranks in the four dimensions are (4, 1, 3, 1), and thus:
[0438]
[0439] Optionally, the score of the picture can be controlled within a certain range, for example, [r mim , r max ], r max is the value of the largest element in R, and r min is the value of the smallest element in R. f i = (f i1 , f i2 , …, f in )T After normalization f' i = (f' i1 ,f' i2 ,…,f' in ) T Then the fusion score is calculated, i.e.
[0440] s i = R T f' i (17)
[0441]
[0442] where f ij is the corrected total number of times the ith picture appears in the lth rank.
[0443] That is, the present scheme is an improvement on the Borda Count to take into account the matching degree and the preset weight corresponding to each dimension
[0444] The present scheme, when calculating the score of each visual media,
[0445] It should be noted that, in order to better protect user privacy and security, meet the principle of user data minimization, and avoid user data reporting to the cloud side as much as possible, the entire search process is completed on the terminal side.
[0446] In addition, the present application provides an electronic device, comprising a memory, a processor and a display, wherein the memory is used to store a program; the processor is coupled with the memory and the display, and is used to execute the program stored in the memory to realize the above-mentioned visual media search method.
[0447] The present application also provides a computer readable storage medium storing a computer program, wherein the computer program can realize one or more steps of the above-mentioned visual media search method when executed by a computer.
[0448] The computer readable storage medium can be a non-transitory computer readable storage medium, for example, the non-transitory computer readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk and an optical data storage device, etc.
[0449] Another embodiment of the present application also provides a computer program product containing instructions. When the computer program product is executed by a computer, it can realize one or more steps of the above-mentioned method.
[0450] The electronic device, the computer readable storage medium, and the computer program product provided in the embodiments are used to execute the corresponding method provided above, and thus the beneficial effects achieved by the electronic device, the computer readable storage medium, and the computer program product can refer to the beneficial effects of the corresponding method provided above, which will not be described here again.
[0451] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other manners. For example, the apparatus embodiments described above are only schematic, and the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another apparatus, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or components shown or discussed can be indirect coupling or communication connection through some interfaces, apparatuses or units, and can be electrical, mechanical or other forms.
[0452] The units described as separate components can or can not be physically separate, and the components shown as units can be one physical unit or a plurality of physical units, that is, can be located in one place or can be distributed to a plurality of different places. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.
[0453] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0454] When the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application essentially or the parts that make contributions to the prior art or all or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium, includes several instructions to make an apparatus (which can be a single chip, a chip, etc.) or a processor execute all or part of the steps of the methods of the embodiments of the present application. The storage medium mentioned above includes: a U disk, a mobile hard disk, a read only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0455] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A visual media search method, suitable for an electronic device, characterized by, The method comprises the following steps: displaying a first interface; the first interface comprises a search box; receiving a search operation of a search statement input into the search box; determining W dimensions to be matched according to the search statement; W is an integer greater than 1; obtaining the matching degree and the matching degree ranking of Y visual media files on the W dimensions to be matched; Y is an integer greater than 1; w is an integer and takes a value from 1 to W; correcting the total number of times that the yth visual media file appears in the uth place according to the matching degree of the yth visual media file on each dimension to be matched in the W dimensions to be matched, to obtain the corrected total number of times that the yth visual media file appears in the uth place; wherein u is an integer and takes a value from 1 to Y; y is an integer and takes a value from 1 to Y; determining the score of the yth visual media according to the corrected total number of times that the yth visual media file appears in each place; determining the search result of the search statement according to the scores of the Y visual media files; displaying the search result.
2. The method of claim 1, wherein, The method for correcting the total number of times that the yth visual media file appears in the uth place according to the matching degree of the yth visual media file on each dimension to be matched in the W dimensions to be matched comprises the following steps: determining a first weight of the wth dimension to be matched for the yth visual media file according to the matching degree of the yth visual media file on the wth dimension to be matched; performing weighted summation on the number of times that the yth visual media file appears in the uth place on each dimension to be matched in the W dimensions to be matched according to the first weights of the W dimensions to be matched for the yth visual media file, to obtain the corrected total number of times that the yth visual media file appears in the uth place.
3. The method of claim 2, wherein, The method for determining a first weight of the wth dimension to be matched for the yth visual media file according to the matching degree of the yth visual media file on the wth dimension to be matched comprises the following steps: obtaining a preset weight of the wth dimension to be matched; correcting the matching degree of the yth visual media file on the wth dimension to be matched according to the preset weight of the wth dimension to be matched, to obtain a corrected matching degree of the yth visual media file on the wth dimension to be matched; determining the corrected matching degree of the yth visual media file on the wth dimension to be matched as the first weight of the wth dimension to be matched for the yth visual media file.
4. The method according to any one of claims 1 to 3, characterized in that, The determination manner of the dimension to be matched comprises multiple manners in the following manners: determining the search statement as the dimension to be matched; determining a semantic subject irrelevant to visual content in the search statement as the dimension to be matched; determining the number of times of user clicks as the dimension to be matched.
5. The method according to any one of claims 1 to 3, characterized in that, The W dimensions to be matched comprise the number of times of user clicks. determining the matching degree of the yth visual media file on the number of times of user clicks according to the number of times of user clicks of the yth visual media file in history.
6. The method according to any one of claims 1 to 3, characterized in that, The W dimensions to be matched comprise the search statement. According to a vector similarity between the visual semantic vector of the yth visual media file and a first sentence semantic vector of the search statement, a matching degree of the yth visual media file on the search statement is determined.
7. The method according to any one of claims 1 to 3, characterized in that, The W dimensions to be matched include a semantic subject irrelevant to visual content in the search statement. According to a matching condition of an attribute of the yth visual media file and the semantic subject irrelevant to visual content, a matching degree of the yth visual media file on the semantic subject irrelevant to visual content in the search statement is determined.
8. The method according to any one of claims 1 to 3, characterized in that, Before the total number of u appearances of the yth visual media file is corrected according to the matching degrees of the yth visual media file on each of the W dimensions to be matched, the method further comprises: The matching degrees of the Y visual media files on the wth dimension to be matched are processed by percentage.
9. The method according to any one of claims 1 to 3, characterized in that, Further comprising: Semantic understanding of the search statement is performed to obtain a first sentence semantic vector; Visual semantic vectors of a plurality of visual media files are obtained; The visual semantic vector of each visual media file is obtained by performing semantic understanding on an image or image frame of the visual media file by using a natural picture understanding model; From the plurality of visual media files, a plurality of first visual media files having visual semantic vectors matching the first sentence semantic vector are determined; The Y visual media files are determined according to the plurality of first visual media files.
10. An electronic device, comprising: Comprise: a memory, a processor and a display, wherein the memory is configured to store a program; the display is configured to display a search page; the processor is coupled with the memory and the display, and is configured to execute the program stored in the memory to implement the method in any one of claims 1 to 9.
11. A computer readable storage medium storing a computer program, characterized in that, The computer program can implement the method in any one of claims 1 to 9 when executed by a computer. The computer program can implement the method in any one of claims 1 to 9 when executed by a computer.
Citation Information
Patent Citations
Method and device for displaying video retrieval results
CN104462573A
Visual media personalized search method and device
CN113641857A