Image library-based image search method and electronic equipment
By introducing interactive methods of sliding gestures and click operations in the gallery application of electronic devices, combined with automatic search function, the problem of inefficient users to find specific content in a large number of image resources is solved, and a more efficient image resource search experience is achieved.
Patent Information
- Application Number
- CN202311589348.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2025-05-30
AI Technical Summary
In the gallery of electronic devices, as the number of pictures and videos increases, the efficiency of users to find specific pictures or videos in a large number of image resources gradually decreases, and there is a lack of convenient and fast search methods.
A gallery-based image search method is provided, guiding the user to enter the search interface by displaying thumbnails of multiple image resources on the user interface of the gallery application and in response to the user's sliding gestures and click operations. In the search interface, users can enter text information to automatically search for image resources and improve search efficiency.
This method significantly improves the efficiency of users in searching pictures or videos in the gallery, provides a convenient and fast search method, and enhances the user experience.
Smart Images

Figure CN120067374A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminals, and in particular, to an image search method and an electronic device based on a gallery. Background Art
[0002] With the continuous improvement of the shooting ability of electronic devices, many users will use electronic devices to take pictures and videos; users will also use electronic devices to download various pictures and videos from the Internet; users will also use electronic devices to receive pictures and videos shared by other devices, etc. The electronic device saves the acquired pictures and videos locally. As the electronic device is used, more and more pictures and videos are saved on the electronic device.
[0003] Generally speaking, the gallery application (hereinafter referred to as the gallery) of the electronic device can provide the pictures and videos saved on the electronic device to the user. The user can find the pictures and videos saved on the electronic device through the user interface provided by the gallery. As time goes by, there are more and more pictures and videos in the gallery, making it difficult to search.
[0004] How to provide a convenient and fast search method for users to help them accurately find the required pictures or videos among a large number of pictures and videos is a problem to be solved. Summary of the Invention
[0005] Embodiments of this application provide an image search method and an electronic device based on a gallery, which can help users more conveniently find the required pictures, videos, etc. in the gallery, and improve the user experience of using the gallery to search for pictures and videos.
[0006] To achieve the above object, the embodiments of this application adopt the following technical solutions:
[0007] In a first aspect, an image search method based on a gallery is provided. The method includes: displaying a first user interface of a gallery application, where the first user interface includes thumbnails of multiple image resources (pictures or videos) in the gallery application; in response to a swipe gesture of the user on the first user interface, the thumbnails of the multiple image resources move positions along the direction of the swipe gesture, and a first button is also displayed on the first user interface; the user can click the first button. In response to the user's click operation on the first button, a search interface including a first search box is displayed, and the first search box is used to receive text information input by the user, so that image resources can be searched in the gallery application according to the text information.
[0008] In this method, the electronic device displays a first user interface of a gallery application, which is used to display pictures and / or videos saved in the gallery application. Upon receiving a sliding gesture from the user on the first user interface (the user may be manually searching for photos or videos), the electronic device displays a first button. The user can easily enter the search interface by clicking the first button, and enter text information in the search box of the search interface to trigger the electronic device to automatically search for image resources in the gallery. The electronic device guides the user to enter the search interface provided by the electronic device conveniently and quickly through the first button, and then the user can enter text information in the search box of the search interface to realize automatic search for image resources, thereby improving the efficiency of the user in finding pictures or videos in the gallery application.
[0009] In combination with the first aspect, in a possible implementation, after the user stops the sliding gesture on the first user interface, a second user interface is displayed, the second user interface includes a second search box. In response to the user clicking the second search box, a search interface is displayed.
[0010] In this method, after the user stops manually searching for photos or videos (the sliding gesture stops), the mobile phone displays a search box, prompting the user to enter text information to realize automatic search. The user conveniently enters the search interface by clicking the search box, and enters text information in the search box of the search interface to trigger the mobile phone to automatically search for image resources in the gallery. This provides a convenient and quick way for users to enter automatic search.
[0011] In a possible implementation, if a next sliding gesture is not detected within a first time period after a sliding gesture is detected, it is determined that the sliding gesture performed by the user on the first user interface stops.
[0012] In combination with the first aspect, in a possible implementation, displaying a first user interface of a gallery application includes: displaying a third user interface of the gallery application, the third user interface including thumbnails of multiple image resources in the gallery application; in response to a sliding gesture of the user on the third user interface, displaying the first user interface of the gallery application; wherein the image resources included in the first user interface are different from the image resources included in the third user interface.
[0013] In this method, in response to an operation by the user to launch the gallery application, a third user interface of the gallery application is displayed, and the interface includes thumbnails of a plurality of image resources (for displaying the image resources). The user can perform a swipe gesture on the third user interface to cause the thumbnails of the image resources to roll in the direction of the swipe gesture to view more image resources. For example, in response to a swipe gesture by the user on the third user interface, the electronic device displays the above-mentioned first user interface, and the first user interface includes a first button. That is to say, during the process of the user manually viewing image resources through a swipe gesture, the electronic device displays the first button to guide the user to quickly enter the automatic search.
[0014] In combination with the first aspect, in a possible implementation manner, the third user interface includes the above-mentioned second search box, and in response to a swipe gesture by the user on the third user interface, the second search box is hidden.
[0015] In this scenario, the third user interface includes a search box. When a swipe gesture by the user on the third user interface is received, it indicates that the user needs to manually search for image resources, and the electronic device hides the second search box.
[0016] In combination with the above possible implementation manner, the method provided in this application can achieve beneficial effects in the following scenarios:
[0017] The electronic device displays a third user interface, and the interface includes a search box. The user can enter the automatic search by clicking on the search box. If the user does not click on the search box but performs a swipe gesture to manually search for the required picture or video, the electronic device hides the search box. During the process of the user manually searching for the required picture or video, the user may find that the manual search efficiency is too low and needs to perform an automatic search again. In the method provided in the embodiments of this application, the electronic device displays a first button to guide the user to quickly enter the automatic search by clicking on the first button, which provides convenience for the user to search for image resources in the gallery.
[0018] In combination with the first aspect, in a possible implementation manner, in response to receiving first text information input by the user in the first search box, a first search result interface is displayed; the first search result interface includes a first thumbnail of a first video, and the first thumbnail includes a first time point; in response to receiving second text information input by the user in the first search box, a second search result interface is displayed; the second search result interface includes a second thumbnail of the first video, and the second thumbnail includes a second time point, and the second time point is later than the first time point.
[0019] In this method, according to the text information input by the user, the electronic device automatically searches for a video corresponding to the text information in the gallery and displays the thumbnail of the video on the search result interface. There may be image frames corresponding to the text information in a video, or there may be image frames not corresponding to the text information; in this method, the thumbnail of the video displays the image and time point of the image frame corresponding to the text information. In this way, the user can conveniently obtain the video corresponding to the text information and accurately know the specific video segment corresponding to the text information in the video.
[0020] Combined with the first aspect, in a possible implementation manner, the image of the first thumbnail is the image frame corresponding to the first time point in the first video, and the image of the second thumbnail is the image frame corresponding to the second time point in the first video.
[0021] Combined with the first aspect, in a possible implementation manner, in response to the user's click operation on the first thumbnail, the first video is played starting from the first time point; in response to the user's click operation on the second thumbnail, the first video is played starting from the second time point.
[0022] In this method, after the user triggers the first thumbnail, the electronic device plays the first video starting from the first time point. After the user triggers the second thumbnail, the electronic device plays the first video starting from the second time point. That is, after the user triggers the first thumbnail or the second thumbnail, the electronic device plays the first video, but the starting time points are different, that is, the image frames shown to the user at the beginning are different. This starting time point is the time point corresponding to the specific video segment corresponding to the text information. That is to say, the electronic device directly shows the user the video segment related to the text information input by the user, without the user having to manually search in the video, improving the user experience.
[0023] Combined with the first aspect, in a possible implementation manner, in response to the user's click operation on the first thumbnail, a first playback interface of the first video is displayed; the first playback interface includes a plurality of marking points, and the marking points are used to indicate the starting position of the video segment corresponding to the first text information in the first video.
[0024] In this method, in a video, the starting moments of all video segments corresponding to the text information input by the user can be marked on the progress bar of the video playback. This enables the user to conveniently know the starting moments of all video segments corresponding to the text information input by the user. Further, the user can click on any marking point to make the mobile phone start playing the video from the position corresponding to the marking point. The user can conveniently view the video segment to be searched.
[0025] In combination with the first aspect, in a possible implementation, the first text information is different from the second text information. The first video includes a first image frame and a second image frame. The first text information matches the first image frame corresponding to the first thumbnail, and the second text information matches the second image frame corresponding to the second thumbnail. The second image frame is after the first image frame.
[0026] Among them, the first text information and the first image frame corresponding to the first thumbnail are matched through the CLIP model, and the second text information and the second image frame corresponding to the second thumbnail are matched through the CLIP model.
[0027] The first text and the first image frame corresponding to the first thumbnail are matched through the CLIP model, that is, the text semantics of the first text match the visual semantics of the first image frame. The second text and the second image frame corresponding to the second thumbnail are matched through the CLIP model, that is, the text semantics of the second text match the visual semantics of the second image frame. In this way, based on the text information input by the user, an image frame matching the text information can be searched. The text semantics of the text information are fully associated with the visual semantics of the image frame, enabling the fusion interaction between the text information and the content of the picture presented by the image frame, thereby improving the accuracy of video search and enhancing the user experience.
[0028] In combination with the first aspect, in a possible implementation, the matching of the first text information and the first image frame corresponding to the first thumbnail through the CLIP model includes: inputting the first text information into the text encoder of the CLIP model to obtain a first text semantic vector; inputting the first image frame into the image encoder of the CLIP model to obtain a first visual semantic vector; based on the first text semantic vector and the first visual semantic vector, matching the first text information and the first image frame; wherein, the vector similarity between the first text semantic vector and the vector of the first clustering center point of the inverted index library is greater than or equal to a first threshold, the first visual semantic vector belongs to the first clustering cluster corresponding to the first clustering center point, the vector similarity between the first text semantic vector and the first visual semantic vector is greater than or equal to a second threshold, and the inverted index library includes clustering clusters corresponding to multiple clustering center points respectively, and the multiple clustering center points are determined by clustering multiple visual semantic vectors in the inverted index library.
[0029] In this way, when the electronic device performs search and matching based on the text semantic vector of the text information input by the user, it can first match it with multiple clustering center points, and then match it with the visual semantic vectors of the clustering clusters of the determined clustering center points, avoiding matching with all the indexes in the index library. Further avoiding search latency, improving video search efficiency, and thus enhancing the user experience.
[0030] In combination with the first aspect, in a possible implementation manner, the first search result interface further includes a third thumbnail of the first video. The first text information matches the third image frame corresponding to the third thumbnail. The first image frame is in the first video segment of the first video, and the third image frame is in the third video segment of the first video. The entity of the first text information matches the entity of the attribute label of the first video segment, and the entity of the first text information matches the entity of the attribute label of the third video segment. The display order of the first thumbnail is before the third thumbnail.
[0031] In this method, the entity of the first text information matches the entity of the attribute label of the first video segment, and the entity of the first text information matches the entity of the attribute label of the third video segment. In this way, based on the text information input by the user, on the basis of searching and matching the text semantic vector, the electronic device also conducts searching and matching of the entities in the text information, that is, the electronic device can perform vector recall and entity recall, and can display video search results that are both vector recall results and entity recall results to the user, further improving the accuracy of video search and enhancing the user experience.
[0032] Among them, the display order of the first thumbnail and the third thumbnail is determined through the following steps:
[0033] Based on the vector similarity between the first visual semantic vector and the first text semantic vector, and the matching degree between the attribute label of the first video segment and the entity of the first text information, determine the first comprehensive matching degree between the first thumbnail and the first text information; based on the vector similarity between the third visual semantic vector (obtained by inputting the third image frame into the image encoder of the CLIP model) and the first text semantic vector, and the matching degree between the attribute label of the third video segment and the entity of the first text information, determine the second comprehensive matching degree between the third thumbnail and the first text information; according to the order from large to small of the comprehensive matching degree, display the first thumbnail before the third thumbnail.
[0034] In this way, the video search results are sorted based on the comprehensive matching degree, ensuring that the video search results ranked in a better position are results that are more matched with the text information input by the user, further enhancing the user experience.
[0035] In combination with the first aspect, in a possible implementation manner, recommendation information is displayed in the first search box; the recommendation information is generated according to at least one of the shooting time, shooting location, person, subject type, and event of the first picture saved in the gallery application.
[0036] In this method, the electronic device displays recommended information to the user in the search box of the gallery. The recommended information is a sentence with natural semantics generated based on the first picture in the gallery. The user can input text information for searching according to the content and format of the recommended information. For example, the user can use the recommended information as the text information in the search box to conduct a search. Since the recommended information is generated based on the first picture in the gallery, the electronic device can search for the first picture and all image resources similar to the content of the first picture in the gallery according to the text information, improving the search hit rate. Moreover, since the recommended information is a sentence with natural semantics, it can more accurately express the purpose of the user's search, narrow the scope of the search results, and help the user more conveniently find the required pictures or videos.
[0037] In combination with the first aspect, in a possible implementation manner, a semantic analysis algorithm is used to perform semantic analysis on the first picture to obtain at least one of the people, subject type, and event of the first picture.
[0038] In combination with the first aspect, in a possible implementation manner, the first picture is a picture taken within a preset time period. For example, the preset time period is more than one month from today.
[0039] In combination with the first aspect, in a possible implementation manner, a prompt message is displayed on the first user interface. The prompt message is used to prompt the user to click the first button to trigger the search for image resources in the gallery application. In this way, the user can know the purpose of the first button according to the prompt message.
[0040] In a second aspect, an electronic device is provided. The electronic device has the function of implementing the method described in the first aspect above. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0041] In a third aspect, an electronic device is provided, including: a processor, a memory, and a display screen; the display screen is used to display the user interface of the gallery application; the memory is used to store computer execution instructions. When the electronic device runs, the processor executes the computer execution instructions stored in the memory so that the electronic device executes the method described in any one of the first aspect above.
[0042] In a fourth aspect, a computer-readable storage medium is provided. Instructions are stored in the computer-readable storage medium. When it runs on a computer, it enables the computer to execute the method described in any one of the first aspect above.
[0043] Among them, for the technical effects brought by any one of the design methods in the second to fourth aspects, reference can be made to the technical effects brought by different design methods in the first aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 FIG. 1 is a schematic diagram of a scenario applicable to the image search method based on a picture library provided in an embodiment of the present application;
[0045] Figure 2 FIG. 2 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;
[0046] Figure 3 FIG. 3 is a schematic diagram of a scenario example of an image search method;
[0047] Figure 4A FIG. 4 is a schematic diagram of a scenario example of the image search method based on a picture library provided in an embodiment of the present application;
[0048] Figure 4B FIG. 5 is a schematic diagram of a scenario example of the image search method based on a picture library provided in an embodiment of the present application;
[0049] Figure 5 FIG. 6 is a schematic diagram of a scenario example of the image search method based on a picture library provided in an embodiment of the present application;
[0050] Figure 6 FIG. 7 is a schematic diagram of a scenario example of the image search method based on a picture library provided in an embodiment of the present application;
[0051] Figure 7 FIG. 8 is a schematic diagram of a scenario example of the image search method based on a picture library provided in an embodiment of the present application;
[0052] Figure 8 FIG. 9 is a schematic diagram of a scenario example of the image search method based on a picture library provided in an embodiment of the present application;
[0053] Figure 9 FIG. 10 is a schematic diagram of a scenario example of the image search method based on a picture library provided in an embodiment of the present application;
[0054] Figure 10 FIG. 11 is a schematic diagram of a scenario example of the image search method based on a picture library provided in an embodiment of the present application;
[0055] Figure 11A FIG. 12 is a schematic diagram of a scenario example of the image search method based on a picture library provided in an embodiment of the present application;
[0056] Figure 11B FIG. 13 is a schematic diagram of a scenario example of the image search method based on a picture library provided in an embodiment of the present application;
[0057] Figure 12 Schematic diagram of an image search method based on an image library provided by an embodiment of the present application;
[0058] Figure 13 Signaling interaction diagram of an image search method based on an image library provided by an embodiment of the present application;
[0059] Figure 14 Schematic diagram of a vector recall process provided by an embodiment of the present application;
[0060] Figure 15 Schematic diagram of the structural composition of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0061] In the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "the", "above-mentioned", "this" and "such" are also intended to include the forms such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of the present application, "at least one" and "one or more" mean one or more than two (including two). The term "and / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist; for example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0062] Reference to "one embodiment" or "some embodiments" etc. described in this specification means that specific features, structures or characteristics described in conjunction with the embodiment are included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways. The term "connection" includes direct connection and indirect connection, unless otherwise stated. "First" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features.
[0063] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0064] Generally speaking, electronic devices such as mobile phones and tablet computers all provide a gallery. Through the user interface in the gallery, users can find the pictures and videos saved on the electronic device. Static pictures usually include a single frame of image, dynamic pictures include multiple frames of images, and videos are also composed of a series of images. For the sake of convenience of description, in the embodiments of the present application, pictures and videos are uniformly referred to as image resources.
[0065] When the number of image resources in the gallery is large, it is relatively slow to locate the pictures or videos needed by the user through manual flipping. For example, the gallery stores thousands of pictures taken by the user with the mobile phone in the past 3 years, and these thousands of pictures are arranged in the order from the most recent shooting time to the oldest. It will be time-consuming and difficult for the user to manually flip through each picture to find a certain picture taken 2 years ago.
[0066] Currently, the gallery generally supports an image search function. The user interface of the gallery includes a search box. After receiving the text information input by the user in the search box, the electronic device searches for the saved image resources according to the text information input by the user, and then displays the image resources corresponding to the text information input by the user to the user. Exemplarily, as Figure 1 shown, taking the electronic device as the mobile phone 100 as an example, the mobile phone 100 displays a desktop interface, and the desktop interface includes application icons of multiple application programs. For example, the desktop interface includes an application icon 101 of the gallery for starting the gallery. Exemplarily, in response to the user's click operation on the application icon 101, the mobile phone 100 starts the gallery and displays the album display interface 102 of the gallery. The album display interface 102 includes a search box 103 and multiple album folders. For example, the multiple album folders may include Figure 1 shown "All Photos", "Camera", "My Favorites", "Screenshots & Screen Recordings", "My Favorites", "Self-created", "Video Editing", etc. The mobile phone 100 receives the text information input by the user in the search box 103, and can search for the image resources that are the same as the text information input by the user in the multiple saved album folders.
[0067] When users use the image search function, how to improve the convenience and accuracy of the search and enhance the user experience is a problem that needs to be solved. The embodiment of the present application provides an image search method based on a gallery, which effectively improves the convenience and accuracy of the user using the image search function in the gallery and enhances the user experience of searching for image resources in the gallery.
[0068] The image search method based on the gallery provided in the embodiment of the present application can be applied to electronic devices including gallery applications. The above electronic devices may include mobile phones, tablet computers, laptops, personal computers (PCs), ultra-mobile personal computers (UMPCs), handheld computers, netbooks, smart home devices (e.g., smart TVs, smart screens, large screens, smart speakers, smart air conditioners, etc.), personal digital assistants (PDAs), wearable devices (e.g., smart watches, smart bracelets, etc.), vehicle-mounted devices, virtual reality devices, etc., and the embodiment of the present application does not impose any restrictions on this.
[0069] In the embodiment of the present application, the electronic device is an electronic device that can run an operating system and install application programs. Optionally, the operating system running on the electronic device can be system, system, System, etc.
[0070] Taking the electronic device as a mobile phone as an example, Figure 2 A schematic diagram of the structure of a mobile phone 100 is shown. The mobile phone 100 may include: a processor 110, a memory 120, an antenna 1, an antenna 2, a mobile communication module 130, a wireless communication module 140, an audio module 150, a camera 160, a display screen 170, a sensor module 180, etc. The sensor module 180 may include a pressure sensor, a fingerprint sensor, a temperature sensor, a touch sensor, etc.
[0071] It is to be understood that the structure shown in this embodiment does not constitute a specific limitation on the mobile phone 100. In other embodiments, the mobile phone 100 may include more or fewer components than shown in the figure, or combine some components, or separate some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0072] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0073] The controller may be the nerve center and command center of the mobile phone 100. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.
[0074] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0075] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0076] It can be understood that the interface connection relationships between the modules illustrated in this embodiment are only illustrative descriptions and do not constitute a structural limitation on the electronic device. In other embodiments, the electronic device may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0077] The wireless communication function of the mobile phone 100 can be implemented through Antenna 1, Antenna 2, Mobile Communication Module 130, Wireless Communication Module 140, Modulation and Demodulation Processor, and Baseband Processor, etc.
[0078] The Mobile Communication Module 130 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the mobile phone 100. In some embodiments, at least some functional modules of the Mobile Communication Module 130 may be provided in the Processor 110. In some embodiments, at least some functional modules of the Mobile Communication Module 130 and at least some modules of the Processor 110 may be provided in the same device.
[0079] The Wireless Communication Module 140 can provide solutions for wireless communications applied to the mobile phone 100, including Wireless Local Area Networks (WLAN) (such as Wi-Fi networks), Bluetooth (BT), Global Navigation Satellite System (GNSS), Frequency Modulation (FM), Near Field Communication (NFC), Infrared (IR), etc.
[0080] In the embodiments of this application, the mobile phone 100 can receive pictures or videos from other devices through the Mobile Communication Module 130 or the Wireless Communication Module 140, and save the received pictures or videos to the picture library.
[0081] The mobile phone 100 realizes the display function through the GPU, the display screen 170, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 170 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The Processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0082] The display screen 170 is used to display images, videos, etc. The display screen 170 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini-LED, a Micro-OLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc.
[0083] In the embodiments of the present application, the display screen 170 can be used to display the user interface of the mobile phone 100, such as the desktop interface, the album display interface of the gallery, etc.
[0084] The mobile phone 100 can implement the shooting function through the ISP, the camera 160, the video codec, the GPU, the display screen 170, and the application processor, etc.
[0085] The ISP is used to process the data fed back by the camera 160. For example, when taking a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera photosensitive element. The optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the noise, brightness, and skin color of the image through algorithms. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 160.
[0086] The camera 160 is used to capture static images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV, etc. format. In some embodiments, the electronic device can include one or N cameras 160, where N is a positive integer greater than 1. In the embodiments of the present application, the static images or videos captured by the camera 160 can be saved to the gallery.
[0087] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0088] The video codec is used to compress or decompress digital videos. The mobile phone 100 can support one or more video codecs. In this way, the mobile phone 100 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0089] The NPU is a neural-network (NN) computing processor. By learning from the structure of biological neural networks, such as the transmission mode between human brain neurons, it can quickly process input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of electronic devices can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.
[0090] The audio module 150 is used to convert digital audio information into an analog audio signal for output, and is also used to convert analog audio input into digital audio signals. The audio module 150 can also be used to encode and decode audio signals. In some embodiments, the audio module 150 can be disposed in the processor 110, or some functional modules of the audio module 150 can be disposed in the processor 110.
[0091] The memory 120 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in the memory 120. For example, in the embodiments of the present application, the processor 110 can execute the instructions stored in the memory 120. The memory 120 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function. The data storage area can store data created during the use of the electronic device (such as picture or video files, etc.). In addition, the memory 120 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0092] The pressure sensor is used to sense pressure signals and can convert the pressure signals into electrical signals. In some embodiments, the pressure sensor may be disposed on the display screen 170. There are many types of pressure sensors, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor may include at least two parallel plates having conductive materials. When a force acts on the pressure sensor, the capacitance between the electrodes changes. The mobile phone 100 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on the display screen 170, the mobile phone 100 detects the intensity of the touch operation according to the pressure sensor. The mobile phone 100 can also calculate the position of the touch according to the detection signal of the pressure sensor.
[0093] The touch sensor, also known as the "touch panel". The touch sensor may be disposed on the display screen 170, and the touch sensor and the display screen 170 form a touch screen, also known as the "touch screen". The touch sensor is used to detect touch operations acting thereon or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 170. In some other embodiments, the touch sensor may also be disposed on the surface of the mobile phone 100, at a different position from that of the display screen 170.
[0094] The mobile phone 100 can detect the operations of the user on the display screen 170 through the pressure sensor or the touch sensor. For example, the operation may be an operation of inputting text information in the search box. For example, the operation is an operation of swiping up or down. For example, the operation is an operation of clicking on a certain control on the user interface.
[0095] Taking the electronic device as a mobile phone as an example, in conjunction with the accompanying drawings, the image search method based on the picture gallery provided in the embodiments of the present application will be described in detail below.
[0096] The mobile phone can display a prompt message in the search box of the picture gallery, and the prompt message is used to prompt the user with the content of the text information to be input into the search box. The user inputs the text information in the search box according to the prompt message. Detecting the operation of the user inputting text information in the search box, the mobile phone can search in the picture gallery according to the text information in the search box to find the image resources corresponding to the text information.
[0097] In one implementation, after the mobile phone obtains a picture or a video, it will set tags for the picture or the video. In one example, the tags may include time (e.g., shooting time, acquisition time, etc.), location (e.g., shooting location, etc.), people (e.g., people's names, people's categories, etc.), categories (such as portrait, birthday, dinner, seaside, road, etc.). A picture or a video may include at least one tag. Some tags may be set by the user, and some tags may be marked by the mobile phone. Exemplarily, when the mobile phone captures a picture or a video through the camera, it can obtain the time and location of capturing the picture or the video, set the time of capturing the picture or the video as the time tag, and set the location of capturing the picture or the video as the location tag. Exemplarily, the user can manually set the name of the person and manually set the category tag as birthday. Exemplarily, if the mobile phone determines that the person in the picture is a boy through image recognition, it will set the category tag of the picture as boy. If the mobile phone determines that the category of the video is a dinner through image analysis, it will set the category tag of the video as dinner.
[0098] For example, the tags of a picture include "Thursday, November 9, 2023", "Beilin District, Xi'an City", "Xiaohua", "Birthday". The tags of another picture include "Friday, July 2, 2021", "Wuhou District, Chengdu City". The tags of yet another video include "Boy", "Girl", "Dinner".
[0099] Currently, the mobile phone gallery supports the user to enter the content of any one tag in the search box. The mobile phone can display information in the search box to prompt the content entered by the user. Exemplarily, such as Figure 3As shown in the figure, the mobile phone 100 displays the album display interface 102 of the photo gallery. The album display interface 102 includes a search box 103, and a prompt message "Photos, People, Locations..." is displayed in the search box 103, which is used to prompt the user that the content entered can be any one of the information such as "Photos", "People", or "Locations". The user can click on the search box 103 and enter text information in the search box according to the prompt message. Exemplarily, in response to the user's click operation on the search box 103, the mobile phone 100 displays a search interface 201. The search interface 201 includes a search box 202, and the user can enter text information in the search box 202. When the mobile phone 100 detects the operation of the user entering text information in the search box 202, it can search in the photo gallery according to the text information in the search box 202. For example, upon receiving "Beijing" entered by the user in the search box 202, the mobile phone 100 searches in the photo gallery for image resources whose location tags are the same as "Beijing City". It should be noted that the mobile phone can perform semantic understanding on the text entered by the user to determine whether the semantics of two texts with different expressions are the same. For example, through semantic understanding, the mobile phone determines that the semantics of "Beijing" and "Beijing City" are the same, that is, whether the user enters "Beijing" or "Beijing City", it can correspond to the location tag "Beijing City". In addition, the location tag being the same as "Beijing" can include location tags such as "Chaoyang District, Beijing City", "Shunyi District, Beijing City", "Xicheng District, Beijing City", etc. Upon receiving "Animals" entered by the user in the search box 202, the mobile phone searches in the photo gallery for image resources whose category tags are the same as "Animals". Upon receiving "2021" entered by the user in the search box 202, the mobile phone searches in the photo gallery for image resources whose time tags are the same as "2021". Exemplarily, as Figure 3 shown in the figure, the user enters the text "Xi'an" in the search box 202, and then clicks on the search box 202. The mobile phone 100 receives "Xi'an" entered by the user in the search box 202 and searches in each album folder in the photo gallery for image resources whose location tags are the same as "Xi'an". As Figure 3 shown in the figure, the mobile phone 100 displays a search result interface 301, and the search result interface 301 is used to display the image resources found in the photo gallery according to the text information in the search box 202. Exemplarily, a total of 1094 image resources are found in each album folder in the photo gallery whose location tags are the same as "Xi'an".
[0100] This method only supports the user to enter simple tags in the search box, usually single tags, such as "2021", "Dinner Party", "Seaside", etc. Generally speaking, the number of image resources found through a single tag is still relatively large, and it is still difficult for the user to find the required pictures or videos among them. For example, Figure 3In the scenario shown, a total of 1,094 image resources with location tags consistent with "Xi'an" were found. It is still not convenient and fast enough for users to manually search for the required pictures or videos among the 1,094 image resources. If the user enters multiple tags in the search box 202, such as "2021" and "Xi'an", since "2021" + "Xi'an" does not match any single tag, the number of image resources searched by the mobile phone 100 may be zero.
[0101] An embodiment of this application provides an image search method based on a picture library, which supports users to input sentences with natural semantics and search for image resources in the picture library according to the sentences with natural semantics. The electronic device displays recommendation information to the user in the search box of the picture library, and this recommendation information is used as an example of the content input by the user to prompt the content and format of the text information input in the search box.
[0102] Still taking the mobile phone 100 as an example, exemplarily, as Figure 4A shown, the mobile phone 100 displays a photo display interface 401. The photo display interface 401 includes a search box 402 and a photo display page 403. The search box 402 is used to trigger the display of an image resource search interface, and the photo display page 403 is used to display thumbnails of image resources (pictures or videos) in the picture library. In one implementation, as Figure 4A shown, the thumbnails of the image resources in the photo display page 403 are arranged in descending order of the time of the image resources.
[0103] In one example, the recommendation information "Try searching for photos of walking by the sea last August" is displayed in the search box 402. This recommendation information is used as an example of the content input by the user. The user can input text information in the search box according to the content and format of this recommendation information to conduct a search.
[0104] Exemplarily, as Figure 4A shown, in response to the user's click operation on the search box 402, the mobile phone 100 displays a search interface 201. The search interface 201 includes a search box 202, and the user can input text information in the search box 202. When the mobile phone 100 detects the operation of the user inputting text information in the search box 202, it can search in the picture library according to the text information in the search box 202. For example, the recommendation information "Try searching for photos of walking by the sea last August" is displayed in the search box 202.
[0105] Among them, the sentence "Taking a walk by the sea last August" has natural semantics. Compared with single tags such as "by the sea" and "dinner", it is easier to hit the image resources that the user really needs to search for. For example, there are 7,230 image resources corresponding to the tag "last year" in the image library, 2,651 image resources corresponding to the tag "August", and 1,985 image resources corresponding to the tag "by the sea". The number of image resources searched using a single tag is relatively large. However, when searching in the image library according to "Taking a walk by the sea last August", the number of image resources found will be significantly reduced.
[0106] Exemplarily, as Figure 4B shown, the user enters the text "Taking a walk by the sea last August" in the search box 202, and then clicks on the search box 202. The mobile phone 100 receives the "Taking a walk by the sea last August" entered by the user in the search box 202 and searches for image resources in the image library that are semantically consistent with "Taking a walk by the sea last August". As Figure 4B shown, the mobile phone 100 displays the search result interface 501, and the search result interface 501 is used to display the image resources found in the image library according to the text information in the search box 202. Exemplarily, a total of 54 image resources that are semantically consistent with "Taking a walk by the sea last August" are found in the image library.
[0107] The image search method based on an image library provided by the embodiments of the present application displays recommended information in the search box of the image library as a sentence with natural semantics. The user can enter text information in the search box according to the content and format of the recommended information, and use a sentence with natural semantics as the input content of the search box. The mobile phone searches for image resources in the image library according to the text information with natural semantics, and can more accurately hit the pictures and videos that the user needs to search for.
[0108] In one implementation, the mobile phone can obtain the constituent elements of the image resources. In one example, the constituent elements include at least one of time, location, person, subject type, and event.
[0109] The constituent element "time" can be obtained according to the shooting time of the image resource; for example, the constituent element "time" includes "Thursday, November 9, 2023", "Friday, July 2, 2021", etc.
[0110] The constituent element "location" can be obtained according to the shooting location of the image resource, such as according to the GPS positioning information when shooting the image resource; for example, the constituent element "location" can include a city (such as Beijing) and / or a scenic spot name (such as the Great Wall).
[0111] The constituent elements "person", "subject type", and "event" can be obtained by performing semantic analysis on image resources. For example, a mobile phone can use a semantic analysis algorithm to perform semantic analysis on images in pictures or videos and generate the content of "person", "subject type", and "event" according to the semantics of the image resources.
[0112] Exemplarily, the constituent element "person" may include:
[0113] Portrait names (such as Xiaohua, Xiaoming, Zhang San, Li Si), acrobats, orchestras, youths, children, infants, etc.
[0114] Exemplarily, the constituent element "subject type" may include:
[0115] Musical instruments, calligraphy, paintings, animals, plants, people, landscapes, beaches, outdoors, buildings, birthdays, etc.
[0116] Exemplarily, the constituent element "event" may include:
[0117] Games, singing, dancing, sports, barbecuing, walking, traveling, etc.
[0118] It can be understood that for an image resource, its corresponding constituent elements may include one or more of time, place, person, subject type, and event. For example, for a picture downloaded from the Internet, if the mobile phone fails to obtain its shooting time, the constituent elements do not include "time" or the content of "time" is empty. For example, for a landscape picture, if the content of "person" is not obtained after semantic analysis, the constituent elements do not include "person" or the content of "person" is empty.
[0119] In one implementation, at least one piece of combined information can be generated by splicing the constituent elements of the image resource according to a preset rule.
[0120] In one example, the mobile phone splices the four constituent elements of the image resource, namely "time", "place", "person", and "event", to generate a piece of combined information.
[0121] In one example, the mobile phone splices three constituent elements of the image resource to generate a piece of combined information. For example, "place", "person", and "event" are spliced, "time", "place", and "person" are spliced, "time", "place", and "event" are spliced, "time", "person", and "event" are spliced, "time", "place", and "subject type" are spliced, etc.
[0122] In one example, the mobile phone splices two constituent elements of the image resource to generate a combined message. For example, "person" and "event" are spliced, "location" and "event" are spliced, "time" and "event" are spliced, "location" and "person" are spliced, "time" and "person" are spliced, "location" and "subject type" are spliced, "time" and "location" are spliced, etc.
[0123] In one example, the mobile phone can also generate a combined message based on one constituent element of the image resource. For example, the constituent element is "time".
[0124] In one implementation, by splicing a fixed splicing word with a combined message, the recommended content corresponding to the image resource can be generated. The fixed splicing words include "try searching", "at", "take a picture of", etc.
[0125] Table 1 shows some examples of the recommended content generated according to different numbers of constituent elements. It should be noted that when generating the recommended content, the content of the constituent elements can be mapped correspondingly. For example, the specific time "Monday, August 7, 2023" is mapped to "today", "the day before yesterday", "August", "last month", "last year", or "this year", etc.
[0126] Table 1
[0127]
[0128]
[0129] The constituent elements of an image resource can generate multiple spliced contents according to the splicing combination methods shown in Table 1. For example, if the constituent elements of the image resource include "person", "time", "location", and "event", 14 different spliced contents can be generated respectively according to the splicing combination methods of numbers 1-14 in Table 1. If the constituent elements of the image resource include "location", "person", and "event", 4 different spliced contents can be generated respectively according to the splicing combination methods of numbers 2, 7, 8, and 10 in Table 1.
[0130] In one implementation, the mobile phone determines any one of the multiple spliced contents as the recommended content corresponding to the image resource.
[0131] In another implementation, each of the above splicing and combination methods corresponds to a priority. In the order from front to back in Table 1, the priorities of the splicing and combination methods decrease in turn. The mobile phone generates the recommended content corresponding to the image resource according to the splicing and combination method with the highest priority among the splicing and combination methods supported by the image resource. Exemplarily, if the constituent elements of an image resource include "location", "person", and "event", but do not include "time", then according to the splicing and combination method with the serial number 2 in Table 1, the fixed splicing words are spliced with "location", "person", and "event" to generate the recommended content corresponding to the image resource.
[0132] In the image search method based on the picture gallery provided by the embodiments of the present application, the recommended information in the search box is generated according to the recommended content corresponding to an image resource in the picture gallery. For example, the recommended content corresponding to an image resource is "try to search for photos of + 'Xiaohua' + 'the day before yesterday' + 'at' + 'Big Wild Goose Pagoda' + 'traveling'", and the recommended information generated according to this recommended content is "try to search for photos of Xiaohua traveling at the Big Wild Goose Pagoda the day before yesterday".
[0133] In one implementation, the mobile phone selects the recommended content corresponding to the first image resource in the picture gallery to generate the recommended information in the search box. In one example, the first image resource is an image resource within a preset time period; for example, the preset time period is more than one month away from today. In one example, the mobile phone updates the first image resource once a day, and the first image resources selected within a preset duration (such as within one week) are not repeated.
[0134] Exemplarily, on the first day, the recommended information displayed in the search box in the picture gallery of the mobile phone is "try to search for photos of Xiaohua traveling at the Big Wild Goose Pagoda the day before yesterday"; on the second day, the recommended information displayed in the search box in the picture gallery of the mobile phone is "try to search for photos of walking by the sea last August"; on the third day, the recommended information displayed in the search box in the picture gallery of the mobile phone is "try to search for photos of animals taken at the zoo";... Within a preset duration (such as within one week), the first image resources selected by the mobile phone every day are not repeated. Correspondingly, the recommended information displayed in the search box in the picture gallery of the mobile phone is not repeated every day.
[0135] It should be noted that the image search method based on the picture gallery provided by the embodiments of the present application can be applied to the search box in any user interface in the picture gallery. Figure 4A Taking the photo display interface 401 including the search box 402 as an example for exemplary illustration does not limit the application scenarios of the embodiments of the present application. Exemplarily, as Figure 5As shown, the mobile phone 100 displays a desktop interface, and the desktop interface includes an application icon 101 of the photo gallery. In response to a user's click operation on the application icon 101, the mobile phone 100 displays an album display interface 102 of the photo gallery, and the album display interface 102 includes a search box 103. A recommended message "Try to search for photos of walking by the sea in August last year" is displayed in the search box 103. This recommended message is used as an example of the text information input by the user. The user can input text information in the search box according to the content and format of this recommended message for searching. Exemplarily, in response to the user's click operation on the search box 103, the mobile phone 100 displays a search interface 201, and the search interface 201 includes a search box 202. The recommended message "Try to search for photos of walking by the sea in August last year" is displayed in the search box 202.
[0136] Reference Figure 4B , the user enters the text "walking by the sea in August last year" in the search box 202, and then clicks on the search box 202. The mobile phone 100 receives the "walking by the sea in August last year" entered by the user in the search box 202 and searches for image resources in the photo gallery that are semantically consistent with "walking by the sea in August last year". As Figure 4B shown, the mobile phone 100 displays a search result interface 501, and the search result interface 501 is used to display the image resources found in the photo gallery according to the text information in the search box 202. Exemplarily, a total of 54 image resources in the photo gallery that are semantically consistent with "walking by the sea in August last year" are found.
[0137] In some embodiments, the mobile phone pre-analyzes the data of the image resources saved in the photo gallery to obtain the recommended content corresponding to each image resource in the photo gallery, and saves the recommended content corresponding to each image resource in the photo gallery. For example, when the processor of the mobile phone is idle, it analyzes the data of the image resources saved in the photo gallery to generate the recommended content corresponding to each image resource in the photo gallery. For example, after each change (such as addition, deletion, editing) of the image resources in the photo gallery, the mobile phone analyzes the data of the image resources saved in the photo gallery to generate the recommended content corresponding to each image resource in the photo gallery. Before the mobile phone displays the search box (for example Figure 5 in the scenario shown, after the mobile phone 100 receives the user's click operation on the application icon 101), according to preset conditions (such as, the shooting time of the image resource is more than one month away from today; such as, the recommended content is not repeated within a week, etc.), it selects a first image resource in the photo gallery, and then obtains the recommended content corresponding to the first image resource from the recommended content saved by the mobile phone. The mobile phone generates a first recommended message according to the recommended content corresponding to the first image resource and displays the first recommended message in the search box.
[0138] In this implementation, the mobile phone generates recommended content corresponding to each image resource in advance during idle time, and searches for recommended content corresponding to a selected image resource when the recommended information needs to be displayed to generate the recommended information. When generating the recommended information, it is not necessary to analyze the image resource to obtain the recommended content, but to directly read the recommended content from the saved information, so that the efficiency of generating the recommended information is high and the delay is small.
[0139] In other embodiments, the mobile phone displays the search box before (for example Figure 5 In the illustrated scenario, after receiving a user's click operation on the application icon 101, the mobile phone 100 selects a first image resource in the gallery according to preset conditions (for example, the image resource is shot more than one month from today; for example, the recommended content is not repeated within a week, etc.), performs data analysis on the first image resource, and generates recommended content corresponding to the first image resource. The mobile phone generates first recommended information according to the recommended content corresponding to the first image resource, and displays the first recommended information in the search box.
[0140] In this implementation, when it is necessary to display recommended information, an image resource is selected, semantic recognition is performed on the selected image resource to obtain recommended content and generate recommended information. There is no need to save a large amount of recommended content, which can save storage space.
[0141] In the image search method based on the gallery provided in the embodiment of the present application, the electronic device displays recommended information to the user in the search box of the gallery, and the recommended information is a sentence with natural semantics generated based on the first image resource in the gallery. The user can enter text information for search based on the content and format of the recommended information. For example, the user can use the recommended information as the text information in the search box to search. Since the recommended information is generated based on the first image resource in the gallery, the electronic device can search for the first image resource and all image resources with similar content to the first image resource in the gallery based on the text information, thereby improving the hit rate of the search. In addition, since the recommended information is a sentence with natural semantics, it can more accurately express the purpose of the user's search, narrow the scope of search results, and help users find the required pictures or videos more conveniently.
[0142] In some scenarios, users may not be able to quickly find the search box to search, but instead manually search for pictures or videos, which is inefficient. The image search method based on the gallery provided in the embodiment of the present application provides a method for quickly entering the search interface, guiding the user to conveniently and quickly enter the search interface provided by the electronic device, and then the user can enter the information in the search box of the search interface to automatically search for image resources.
[0143] Still take the electronic device displaying a photo display interface as an example for description. For example, Figure 6As shown, the mobile phone 100 displays a photo display interface 401. The photo display interface 401 includes a search box 402 and a photo display page 403. The search box 402 is used to trigger the display of an image resource search interface, and the photo display page 403 is used to display thumbnails of image resources (pictures or videos) in the gallery. The user may swipe up or down on the photo display page 403. In one example, when an upward swipe gesture from the user is received on the photo display page 403, in response to the upward swipe gesture of the user, the mobile phone 100 displays a photo display interface 404, and the search box is no longer displayed on this photo display interface 404. Exemplarily, the photo display interface 404 includes a photo display page 405, and the photo display page 405 is used to display thumbnails of image resources (pictures or videos) in the gallery. In one implementation, the user continuously makes an upward swipe gesture on the photo display page 405. In response to the upward swipe gesture of the user on the photo display page 405, the thumbnails of the image resources displayed within the photo display page 405 scroll upward accordingly. Exemplarily, as Figure 6 shown, the thumbnails of the image resources in the photo display page 403 are arranged in descending order of the time of the image resources. What is displayed within the photo display page 403 are the image resources taken on November 9, 2023 and the image resources taken on October 30, 2023. In response to the upward swipe gesture of the user, what is displayed within the photo display page 405 are the image resources taken on October 30, 2023 and the image resources taken on September 17, 2023.
[0144] Continue to refer to Figure 6 , in some embodiments, if it is detected that the upward swipe gesture within the photo display page 405 has not stopped, that is, during the process that the user continuously makes an upward swipe gesture within the photo display page 405, the mobile phone 100 displays a prompt button 406, and a prompt message "Find photos, try the search function" is displayed on the prompt button 406. The prompt message displayed on the prompt button 406 is used to prompt the user to click the prompt button 406 to trigger the display of an image resource search interface. Optionally, the mobile phone 100 displays this prompt button 406 within the photo display page 405. It should be noted that the continuously making an upward swipe gesture described in the embodiments of the present application means that the mobile phone detects the next upward swipe gesture within a certain duration (such as 0.5 seconds) after detecting the end of one upward swipe gesture. If no next upward swipe gesture is detected within a certain duration (such as 0.5 seconds) after the end of one upward swipe gesture, it is determined that the upward swipe gesture has stopped.
[0145] The user can click the prompt button 406 to trigger the display of an image resource search interface. Exemplarily, as Figure 6As shown, in response to the user's click operation on the prompt button 406, the mobile phone 100 displays a search interface 201. The search interface 201 includes a search box 202, where the user can enter text information. When the mobile phone 100 detects the operation of the user entering text information in the search box 202, it can search for corresponding image resources in the photo gallery according to the text information in the search box 202.
[0146] In this way, during the process of the user manually searching for photos or videos, the user can conveniently enter the search interface according to the prompt information by clicking a button, enter text information in the search box of the search interface, and trigger the mobile phone to automatically search for image resources in the photo gallery.
[0147] In one implementation, recommended information is displayed in the search box of the search interface. This recommended information is information with natural semantics generated by the mobile phone according to the recommended content corresponding to the first image resource in the photo gallery. Exemplarily, as Figure 6 shown, the recommended information "Try searching for photos of walking by the sea in August last year" is displayed in the search box 202. When the mobile phone 100 receives the input of "walking by the sea in August last year" in the search box 202, it will search for image resources with the same semantics as "walking by the sea in August last year" in the photo gallery.
[0148] Reference Figure 7 , in some embodiments, the mobile phone 100 determines that the upward sliding gesture in the photo display page 405 stops. In one example, the mobile phone 100 does not detect an upward sliding gesture within a certain time period (such as 0.5 seconds), and determines that the upward sliding gesture in the display page 405 stops. The thumbnail of the image resource displayed in the photo display page 405 stops scrolling upward. Exemplarily, as Figure 7 shown, the mobile phone 100 displays a photo display interface 407. The photo display interface 407 includes a search box 408 and a photo display page 409. The search box 408 is used to trigger the display of the image resource search interface, and the photo display page 409 is used to display thumbnails of image resources (pictures or videos) in the photo gallery.
[0149] In one implementation, recommended information is displayed in the search box 408. This recommended information is information with natural semantics generated by the mobile phone according to the recommended content corresponding to the first image resource in the photo gallery. Exemplarily, as Figure 7 shown, the recommended information "Try searching for photos of walking by the sea in August last year" is displayed in the search box 408.
[0150] The user can click on the search box 408 to trigger the display of the image resource search interface. Exemplarily, as Figure 7As shown, in response to the user's click operation on the search box 408, the mobile phone 100 displays a search interface 201. The search interface 201 includes a search box 202 where the user can enter text information. When the mobile phone 100 detects the operation of the user entering text information in the search box 202, it can search for corresponding image resources in the picture gallery according to the text information in the search box 202.
[0151] In this way, after the user stops manually searching for photos or videos, the mobile phone displays a search box, prompting the user that they can enter text information to achieve automatic search. The user can easily enter the search interface by clicking on the search box, and enter text information in the search box of the search interface, triggering the mobile phone to automatically search for image resources in the picture gallery.
[0152] In one implementation, recommended information is displayed in the search box of the search interface. This recommended information is information with natural semantics generated by the mobile phone according to the recommended content corresponding to the first image resources in the picture gallery. Exemplarily, as Figure 7 shown, the recommended information "Try searching for photos of walking by the sea in August last year" is displayed in the search box 202. When the mobile phone 100 receives the input "walking by the sea in August last year" entered by the user in the search box 202, it will search for image resources in the picture gallery that are semantically consistent with "walking by the sea in August last year".
[0153] It should be noted that the embodiments of the present application are introduced by taking the example that the mobile phone detects the user's upward sliding gesture in the photo display interface. This example should not constitute a limitation on the applicable scenarios of the embodiments of the present application. In some other embodiments, the applicable scenarios of the embodiments of the present application also include: the mobile phone detects an upward sliding gesture on other pages in the picture gallery for displaying thumbnail images of image resources to the user, or detects a downward sliding gesture on other pages for displaying thumbnail images of image resources to the user. The specific implementation methods in these scenarios can refer to Figure 6 and Figure 7 the implementation method of detecting an upward sliding gesture shown in the photo display interface. The implementation methods are similar, and will not be shown one by one in the embodiments of the present application.
[0154] The image search method based on the picture gallery provided by the embodiments of the present application displays a prompt button and prompt information when the user slides up or down to search for pictures or videos, prompting the user to trigger the mobile phone to display the search interface, so that the mobile phone automatically searches for image resources according to the input text information of the user in the search box. After the user finishes sliding up or down to search for pictures or videos, the mobile phone displays a search box and recommended information, prompting the user to trigger the mobile phone to display the search interface, so that the mobile phone automatically searches for image resources according to the input text information of the user in the search box. The mobile phone prompts the user to use the automatic image search function of the mobile phone more conveniently through various methods, improving the convenience of the user to search for image resources in the picture gallery.
[0155] The mobile phone receives the text information input by the user in the search box on the gallery user interface, and searches for the image resources corresponding to the text information according to the text information. These image resources can include pictures and can also include videos.
[0156] In one implementation, a text analysis algorithm is used to semantically understand the text information input by the user, and the constituent elements corresponding to the text information are obtained, including at least one of time, place, person, subject type, and event.
[0157] For pictures, the shooting time and shooting location of the pictures can be obtained according to the picture information saved in the mobile phone. A semantic analysis algorithm can also be used to perform semantic analysis on the images in the pictures, and one or more of the person, subject type, and event included in the pictures can be obtained according to the semantics of the pictures. In this way, the constituent elements of the pictures can be obtained.
[0158] If it is determined that the constituent elements included in the text information are consistent with the constituent elements included in the picture, it is determined that the picture is the image resource corresponding to the text information.
[0159] It can be understood that a video is composed of some image frames. In one implementation, if it is determined that a video includes N image frames corresponding to the text information, it is determined that the video is the image resource corresponding to the text information. Wherein, N is a preset value. For example, N = 1, N = 24, etc.
[0160] The mobile phone determines the pictures and videos corresponding to the text information according to the text information, and can then display the search results to the user. For example, the thumbnails of these pictures and videos are displayed in the user interface of the gallery.
[0161] Exemplarily, as Figure 8 shown, the mobile phone 100 displays a search interface 201, and the search interface 201 includes a search box 202. The user inputs the text "taking a walk by the sea in August last year" in the search box 202, and then clicks on the search box 202. The mobile phone 100 receives the "taking a walk by the sea in August last year" input by the user in the search box 202, and searches for the image resources with the same semantics as "taking a walk by the sea in August last year" in the gallery. As Figure 8As shown, the mobile phone 100 displays a search result interface 501, and the search result interface 501 is used to display the image resources found in the picture gallery according to the text information in the search box 202. Exemplarily, a total of 54 image resources with the same semantics as "taking a walk by the sea in August last year" are found in the picture gallery. Optionally, in one example, the search result interface 501 includes a page 502, a page 503, and a page 504; wherein, the page 502 is used to display pictures and videos corresponding to the text information input by the user, the page 503 is used to display pictures corresponding to the text information input by the user, and the page 504 is used to display videos corresponding to the text information input by the user.
[0162] In some embodiments, the video corresponding to the text information includes image frames corresponding to the text information and image frames not corresponding to the text information. Exemplarily, as Figure 9 shown, Video A is the video corresponding to "taking a walk by the sea in August last year" searched by the mobile phone 100. Video A includes many image frames. Among them, the image frames from 00:00 to 02:17 do not correspond to "taking a walk by the sea in August last year". For example, the image frames from 00:00 to 02:17 are the image frames corresponding to "playing volleyball by the sea in August last year"; the image frames from 02:18 to 06:21 are the image frames corresponding to "taking a walk by the sea in August last year"; the image frames from 06:22 to 08:32 do not correspond to "taking a walk by the sea in August last year". For example, the image frames from 06:22 to 08:32 are the image frames corresponding to "running by the sea in August last year".
[0163] In one implementation, the video thumbnail image displayed by the mobile phone in the search result interface is the first image frame in the video, and the first image frame is one of the image frames corresponding to the text information.
[0164] In one example, the first image frame is any one of the image frames corresponding to the text information.
[0165] In another example, the first image frame is the first frame among the image frames corresponding to the text information. In one implementation, the mobile phone judges frame by frame in the time order of the image frames in the video from the front to the back. When the first frame image frame corresponding to the text information is found, the first frame image frame corresponding to the text information is determined as the first image frame. By judging frame by frame, the first frame image frame corresponding to the text information can be accurately determined.
[0166] In another example, each video is divided into one or more video segments according to a preset condition; for example, the preset condition is to divide a video segment every 0.5 seconds according to the playing duration, or divide a video segment every M (for example, M = 12) frames. Each video segment includes a representative frame, which can be used to represent the video segment. For example, the representative frame is the first frame in a video segment, or the last frame in a video segment, or the frame with the highest pixels in a video segment, or the optimal frame in a video segment, or any frame in a video segment, etc. If the representative frame of a video segment is an image frame corresponding to the text information, then it is determined that the video segment is the video segment corresponding to the text information. In one implementation, the first image frame is a frame image in the first video segment corresponding to the text information (for example, this frame image is the representative frame of this video segment). For example, video segments 276 to 762 in video A are video segments corresponding to the text information. The representative frame of the 276th video segment in video A is the first image frame. For example, the representative frame of the 276th video segment can be the first frame of the 276th video segment, or the last frame of the 276th video segment, or the frame with the highest pixels in the 276th video segment, etc.
[0167] For example, as Figure 8 shown, the thumbnail 5021 shown on page 502 is the thumbnail of video A corresponding to "taking a walk by the sea last August", and the image of the thumbnail 5021 is the first image frame in video A.
[0168] For example, as Figure 8 shown, the thumbnail 5041 shown on page 504 is the thumbnail of video A corresponding to "taking a walk by the sea last August", and the image of the thumbnail 5041 is the first image frame in video A.
[0169] In some embodiments, the mobile phone displays the position of the first image frame in the video on the search result interface; for example, the position of the first image frame in the video can be reflected by the playing time corresponding to the first image frame and the total playing duration of the video. In this way, the user can conveniently find the content to be searched without manually searching for the required image frame in the video.
[0170] For example, the playing time of the first image frame is 02 minutes and 18 seconds, and the total playing duration of video A is 08 minutes and 32 seconds. As Figure 8As shown, "02:18 / 08:32" is superimposed on the image of the thumbnail 5021, indicating that the image of the thumbnail 5021 is a frame at the 02nd minute and 18th second in a video with a total playing duration of 08 minutes and 32 seconds. The corresponding position (next to the thumbnail 5041) of the thumbnail 5041 shows "02:18 / 08:32", indicating that the image of the thumbnail 5041 is a frame at the 02nd minute and 18th second in a video with a total playing duration of 08 minutes and 32 seconds.
[0171] In the image search method based on a gallery provided by an embodiment of the present application, the position of the first image frame in the video is displayed in the search result interface, so that the user can conveniently locate to the position corresponding to the text information input by the user in the video. Exemplarily, Figure 8 in which the text information input by the user is "taking a walk by the sea last August", the playing time of the first image frame corresponding to "taking a walk by the sea last August" is the 02nd minute and 18th second, and the time point (02 minutes and 18 seconds) is displayed on the thumbnail of video A shown in the search result interface, and the image of the thumbnail of video A shown in the search result interface is the image frame (the first image frame) corresponding to the 02nd minute and 18th second. In another example, as Figure 10 shown, the text information input by the user in the input box 202 is "running by the sea last August", and the playing time of the first image frame corresponding to "running by the sea last August" is the 06th minute and 21st second. In response to receiving the user input of "running by the sea last August", the mobile phone 100 displays a search result interface 505, and the search result interface 505 is used to display the image resources found in the gallery according to the text information in the search box 202. Exemplarily, a total of 68 image resources with the same semantics as "running by the sea last August" are found in the gallery. Optionally, in one example, the search result interface 505 includes a page 506, a page 507, and a page 508; wherein, the page 506 is used to display pictures and videos corresponding to the text information input by the user, the page 507 is used to display pictures corresponding to the text information input by the user, and the page 508 is used to display videos corresponding to the text information input by the user. Among them, the thumbnail 5061 displayed in the page 506 is the thumbnail of video A corresponding to "running by the sea last August", and the image of the thumbnail 5061 is the first image frame determined according to "running by the sea last August" in video A. The thumbnail 5081 displayed in the page 508 is the thumbnail of video A corresponding to "running by the sea last August", and the image of the thumbnail 5081 is the first image frame determined according to "running by the sea last August" in video A. Exemplarily, the playing time of the first image frame corresponding to "running by the sea last August" in video A is the 06th minute and 21st second, and the total playing duration of video A is 08 minutes and 32 seconds. As Figure 10As shown, "06:21 / 08:32" is superimposed on the image of thumbnail 5061, indicating that the image of thumbnail 5061 is a frame at the 06th minute and 21st second in a video with a total playing duration of 08 minutes and 32 seconds. The corresponding position (next to thumbnail 5081) of thumbnail 5081 shows "06:21 / 08:32", indicating that the image of thumbnail 5081 is a frame at the 06th minute and 21st second in a video with a total playing duration of 08 minutes and 32 seconds.
[0172] In some embodiments, in response to a user's click operation on any video thumbnail in the search result interface, the mobile phone plays the video corresponding to the video thumbnail. In one implementation, the mobile phone starts playing the video corresponding to the video thumbnail from the position of the image frame (the first image frame) corresponding to the image of the video thumbnail.
[0173] Exemplarily, as Figure 11A shown, the mobile phone 100 displays a search result interface 501, and the search result interface 501 is used to display image resources in the gallery corresponding to the text information "taking a walk by the sea last August" input by the user. The search result interface 501 includes a page 502, a page 503, and a page 504. The page 502 includes a thumbnail 5021 of video A, and the image of the thumbnail 5021 is the first image frame in video A. The playing time of the first image frame is the 02nd minute and 18th second, and the total playing duration of video A is 08 minutes and 32 seconds.
[0174] In response to the user's click operation on the thumbnail 5021, the mobile phone 100 plays video A and displays a playing interface 601 of video A. In one implementation, the mobile phone starts playing video A from the first image frame. Exemplarily, as Figure 11A shown, in response to the user's click operation on the thumbnail 5021, the mobile phone 100 displays a playing interface 601 of video A. The playing interface 601 includes a progress bar 602. Exemplarily, the progress bar 602 is directly positioned at the 02nd minute and 18th second, and the mobile phone 100 starts playing video A from the 02nd minute and 18th second.
[0175] In this method, the mobile phone starts playing the video from the image frame (the first image frame) corresponding to the text information input by the user, which can directly locate the content that the user needs to find, is convenient and fast, and improves the user experience.
[0176] In some embodiments, a video includes multiple video segments corresponding to the text information input by the user. Exemplarily, as Figure 11B shown, video A includes two video segments corresponding to the text information "taking a walk by the sea last August"; the first video segment includes the image frames from the 02nd minute and 18th second to the 06th minute and 21st second, and the second video segment includes the image frames from the 07th minute and 03rd second to the 08th minute and 32nd second.
[0177] Correspondingly, the start positions of each video segment corresponding to the text information input by the user are marked on the progress bar 602 in the playback interface 601. Exemplarily, as Figure 11B shown, the progress bar 602 includes marking points 603 and 604. The position indicated by the marking point 603 is the starting moment of the first video segment in video A corresponding to "taking a walk by the sea in August last year", and the position indicated by the marking point 604 is the starting moment of the second video segment in video A corresponding to "taking a walk by the sea in August last year". In this way, the user can clearly see the starting moments of all video segments in video A corresponding to the text information input by the user at a glance.
[0178] In one implementation, the mobile phone starts playing video A from the starting moment of the first video segment in video A corresponding to the text information input by the user. For example, as Figure 11B shown, the mobile phone 100 starts playing video A from the position indicated by the marking point 603.
[0179] In one implementation, when the mobile phone receives a click operation on any marking point on the progress bar 602, the mobile phone 100 starts playing video A from the position indicated by this marking point. For example, referring to Figure 11B , if the mobile phone 100 receives a click operation on the marking point 604, the mobile phone 100 starts playing video A from the position indicated by the marking point 604.
[0180] In this method, in a video, the start positions of all video segments corresponding to the text information input by the user can be marked on the progress bar during video playback. This enables the user to conveniently know the starting moments of all video segments corresponding to the text information input by the user. Further, the user can click on any marking point to make the mobile phone start playing the video from the position corresponding to this marking point. The user can conveniently view the video segments that need to be searched.
[0181] In the method for image search based on a picture gallery provided in the embodiments of the present application, the mobile phone receives the text information input by the user in the input box and searches for the corresponding image resources in the picture gallery according to the input text information. Next, a specific implementation method for searching for image resources in the picture gallery according to the input text information will be introduced in conjunction with the accompanying drawings.
[0182] To enable those skilled in the art to more clearly understand the technical solutions of the embodiments of the present application, the vocabulary involved in the embodiments of the present application will be described first. It can be understood that this description is for a clearer understanding of the embodiments of the present application and does not necessarily constitute a limitation on the embodiments of the present application.
[0183] Video frame: Also known as an image frame, it refers to any frame of a video. A frame is a still image in a video, and consecutive frames can form a video.
[0184] Video segment: It refers to the segments obtained by dividing a video. In some embodiments, the video can be first frame-split to obtain individual video frames of the video, and then a video segmentation algorithm can be used to segment the video including multiple video frames to obtain video segments. For the specific implementation method, refer to the introduction of the following embodiments.
[0185] Representative frame: It is a video frame in a video segment, which can be used to represent the video segment. Exemplarily, the representative frame can be the starting frame, ending frame, a randomly selected video frame, or the optimal frame with the highest score, etc.
[0186] CLIP model: The CLIP (contrastive language-image pre-training) model is a pre-trained neural network model for matching image and text information. In some embodiments, based on the text encoder and image encoder of the CLIP model for contrastive learning, a text encoder for outputting the text semantic vector of text information and an image encoder for outputting the visual semantic vector of a picture or video frame can be trained.
[0187] Text semantic vector: It can be obtained by inputting text information into a text encoder, and it is a vector that can characterize the semantic features of the entire text information. Exemplarily, the text encoder can adopt models such as Transformer commonly used in natural language processing (NLP), which is not limited in the embodiments of this application. In the embodiments of this application, the text information input by the user in the text box can be input into the text encoder to obtain the text semantic vector of the text information.
[0188] Visual semantic vector: It can be obtained by inputting a picture or a video frame into an image encoder. Exemplarily, the image encoder can adopt a CNN model or a VIT model, which is not limited in this application. In the embodiments of this application, the representative frame of the video segment can be input into the image encoder to obtain the visual semantic vector of the representative frame.
[0189] Vector similarity: It is used to describe the similarity degree between two vectors (for example, the text semantic vector and the visual semantic vector). In the embodiments of this application, by comparing the similarity between the text semantic vector and the visual semantic vector, the picture or video frame that matches the text information can be determined. Exemplarily, the vector similarity can be calculated by the cosine similarity calculation formula. Of course, it can also be calculated by other methods.
[0190] Entity: A word with a specific meaning in text information. Exemplarily, an entity may include, but is not limited to, time, location, person name, organization name, proper noun, etc. in text information. In some embodiments, named entity recognition (NER) technology may be used to identify entities with specific meanings in text information, and this application does not limit this.
[0191] Image parameter: Used to indicate the display characteristics of a picture or video frame. Exemplarily, image parameters may include the jitter degree, clarity, pixel value, etc. of a video frame, and this application does not limit this.
[0192] Attribute label: Information used to indicate the attributes of a video or video segment. In the embodiments of this application, the attribute labels of a video may include the video shooting location, video shooting time, person, subject type, or event, etc.
[0193] In some embodiments, a video includes multiple video segments, and the text information input by the user may match one or more video segments of a video.
[0194] As introduced above, a representative frame can be used to represent the corresponding video segment. Therefore, the visual semantic vector of the representative frame can represent the visual semantic vectors of multiple video frames in the video segment, that is, the visual semantic vector of the representative frame is fully associated with the video segment. When the user inputs complex text information based on the picture content shown in the video, this application matches the text semantic vector corresponding to the text information with the visual semantic vector of the representative frame, which can realize the fusion interaction between the text information and the picture content shown in the video segment, improve the accuracy of video search, and thus enhance the user experience.
[0195] It should be understood that if the vector similarity between the visual semantic vector of the representative frame of a video segment and the text semantic vector of a certain text information exceeds the preset vector similarity threshold, when the user inputs this text information, this video segment can be searched.
[0196] In some embodiments, the mobile phone can perform frame splitting on the video to obtain individual frames in the charging or screen-off state, use the video segment algorithm to segment the video based on multiple video frames to obtain video segments, then determine the representative frame of the video segment and the visual semantic vector of the representative frame, and subsequently construct the attribute label of the video segment based on the visual semantic vector of the representative frame. The specific implementation method can be seen in the detailed introduction in the following embodiments. It should be understood that the pictures stored in the picture library do not need to be frame-split, and the attribute labels of the pictures can be constructed by referring to the method of constructing the attribute labels of video segments based on the representative frames of video segments.
[0197] In one implementation, a mobile phone may include functional modules such as a gallery service module, a search module, a multimodal understanding module, and a natural language understanding module.
[0198] Among them, the gallery service module stores pictures / videos obtained under operations such as user shooting, downloading, screenshotting, or screen recording; it is also used to store representative frames of video segments, visual semantic vectors of representative frames, and other information; it also receives text information input by the user so that the mobile phone can perform search matching of videos based on the text information input by the user; it also displays pictures / videos obtained under operations such as user shooting, downloading, screenshotting, or screen recording.
[0199] The search module is used to construct attribute tags corresponding to videos based on the information stored in the above gallery service module, and is also used to perform vector recall among multiple attribute tags based on the text semantic vector corresponding to the text information, and perform entity recall among multiple attribute tags based on the entities in the text information. It is also used to sort the vector recall results and entity recall results to obtain search results.
[0200] The multimodal understanding module is used to perform frame splitting on the video to obtain individual video frames, and then use a video segmentation algorithm to perform segmentation processing on the video to obtain video segments. Subsequently, the representative frames of the video segments and the visual semantic vectors of the representative frames are determined.
[0201] The natural language understanding module is used to identify entities in the text information input by the user.
[0202] The image search method based on a gallery provided by the embodiments of the present application can be implemented based on the interaction among modules such as the gallery service module, the search module, the multimodal understanding module, and the natural language understanding module.
[0203] The following combines Figure 12 to introduce in detail the image search method based on a gallery provided by the present application. As Figure 12 shown, the image search method based on a gallery includes an index construction stage and a search stage.
[0204] In the index construction stage, in some embodiments, an index can be constructed for each video frame of the video. In some embodiments, if this solution is implemented on a terminal device such as a mobile phone, constructing an index for each video frame of the video will increase the latency required for searching and reduce the user experience during the search stage. Optionally, in some embodiments, the video can be segmented, and a video index can be constructed in units of segments.
[0205] Exemplarily, as Figure 12As shown in the figure, first, the video is frame-split to obtain multiple video frames of the video; then, a video segmentation algorithm is used to segment the video based on the multiple video frames of the video to obtain multiple video segments such as video segment 1 and video segment 2; subsequently, representative frames can be selected from the multiple video frames in the video segment, and the multiple video frames in the video segment are respectively evaluated for scores, and the video frame with the highest score in a video segment is determined as the representative frame of the video segment; the representative frame is input into the image encoder of the CLIP model, and the visual semantic vector of the representative frame is output; an index is constructed based on the visual semantic vector corresponding to the representative frame, and an inverted index library is obtained based on the constructed index.
[0206] It should be noted that the above method for selecting representative frames is only an example, and the starting frame, middle frame, ending frame or a random video frame of the video segment can also be selected as the representative frame. In the search stage, the user inputs text information in the search box; the text information is input into the text encoder of the CLIP model, and the text semantic vector corresponding to the text information is output; vector recall is performed from the inverted index library based on the text semantic vector corresponding to the text information; the text information is sent to the natural language understanding module; the natural language understanding module performs entity recognition on the text information to obtain the entities in the text information; entity recall is performed from the inverted index library based on the entities in the text information; based on vector recall and entity recall, the vector recall result and the entity recall result are returned from the inverted index library; the vector recall result and the search result are sorted to obtain the search result; the search result is displayed on the search result interface.
[0207] In some embodiments, the image encoder of the CLIP model and the text encoder of the CLIP model can be trained by a training method based on contrastive learning.
[0208] Next Figure 13 Still taking the electronic device as the mobile phone 100 as an example, the image search method based on the picture library provided by the embodiments of the present application will be described in detail in combination with the picture library service module, the search module, the multimodal understanding module and the natural language understanding module.
[0209] As Figure 13 shown, the image search method based on the picture library provided by the embodiments of the present application can be divided into the following two stages: the index construction stage and the search stage.
[0210] First, Figure 13 The steps included in the index construction stage will be introduced in detail.
[0211] S501: The picture library service module receives the user's new operation or modification operation on the video.
[0212] The new operation of a video refers to the operation in which a user stores a video in the gallery application. Exemplarily, the new operation of a user for a video can be an operation in which the user shoots a video through the mobile phone 100, downloads a video, or records the screen of the mobile phone 100. The modification operation of a video refers to the operation in which a user modifies a video that has been stored in the gallery application. Exemplarily, the modification operation of a user for a video can be an operation in which the user crops, splices, adds special effects, or adds subtitles to the video, etc.
[0213] In some embodiments, the gallery service module receiving the new operation or modification operation of a user for a video includes: the gallery service module receiving the new operation or modification operation of the user for the attribute tags of the video.
[0214] Exemplarily, the attribute tags of a video can include the video acquisition location (such as the location where the video is shot, the download source of the video, etc.), the video acquisition time (such as the shooting time, the download time, or the screen recording time), the person name newly added or modified by the user for the video, etc.
[0215] S502: The gallery service module stores the video.
[0216] In response to the new operation or modification operation of a user for a video, the gallery service module stores the video in the mobile phone 100. Exemplarily, the video can be stored in the local files of the mobile phone 100, and the user can view the video through various channels such as the local folder of the mobile phone 100, the gallery application, etc. Another exemplarily, with the authorization operation of the user, the video can be stored in the cloud for backup to relieve the memory pressure of the mobile phone 100.
[0217] In some embodiments, in response to the new operation or modification operation of the user for the attribute tags of the video, the gallery service module can store the attribute tags of the video.
[0218] S503: The gallery service module calls the multimodal understanding module to determine the representative frame of the video.
[0219] It can be understood that a video usually includes multiple video frames. Assuming that the corresponding visual semantic vectors are determined for each video frame and stored as indexes, that is, a video corresponds to a large number of indexes, which will waste a large amount of computing resources and a large amount of storage space. At the same time, there may be noise information in a large number of indexes, and it also takes time to perform search matching, which is likely to affect the video search results.
[0220] Therefore, in the embodiments of the present application, the multi-modal understanding module first performs segmentation processing on the video, and determines the corresponding representative frames from each video segment, so as to determine the corresponding visual semantic vectors for the representative frames and store them as indexes in the subsequent steps, greatly reducing the number of indexes corresponding to the video. In this way, the representative frames corresponding to multiple video segments can more completely represent the video semantics of the video, and at the same time, costs can be saved.
[0221] In some embodiments, the gallery service module may call the computer vision service provided by the multi-modal understanding module to determine the representative frames of the video. The computer vision service refers to that the multi-modal understanding module performs video semantic understanding on the video, and then the multi-modal understanding module determines the representative frames of the video.
[0222] Video semantic understanding means enabling the mobile phone to understand the meaning expressed by the content shown in the video, such as understanding information such as the types, quantities, positions of objects in the video, and the relationships between objects.
[0223] In some embodiments, the computer vision service provides services such as frame splitting processing of the video, segmenting the video using a video segmentation algorithm, and determining the representative frames of the video segments. Then the process of determining the representative frames of the video can be mainly broken down into the following steps 1-step 3:
[0224] Step 1: The multi-modal understanding module uses the computer vision service to perform frame splitting processing on the video.
[0225] It can be understood that a video is composed of multiple video frames, and each video frame is a still picture in the video, that is, each video frame can be regarded as an image. Frame splitting processing means that the multi-modal understanding module uses the computer vision service to decompose the video into individual video frames.
[0226] In some embodiments, during the process of the multi-modal understanding module performing frame splitting processing on the video, frame identifiers and time points can be marked for each obtained video frame. A frame identifier is used to uniquely mark a video frame, and different video frames can be distinguished based on the frame identifier. The time point refers to the time when the video frame appears in the video.
[0227] In some embodiments, during the process of the multi-modal understanding module using the computer vision service to perform frame splitting processing on the video, the main type of each video frame can be recognized.
[0228] Step 2: The multi-modal understanding module uses the computer vision service to perform segmentation processing on the video.
[0229] Segmentation processing means using the video segmentation algorithm provided by the computer vision service to divide the video into multiple video segments.
[0230] It can be understood that when a video is played, the content it presents changes continuously as the video frames are played in sequence, but the degree of change varies. Therefore, similar video frames (with a small degree of change) can be classified into a video segment.
[0231] In some embodiments, the degree of change between a video frame and its adjacent video frames can be measured based on the main type of the video frame and the image parameters of the video frame, and then the video can be segmented. Exemplarily, the image parameters of the video frame can include the jitter degree, clarity, pixel value, etc. of the video frame, and the present application does not limit this.
[0232] Among them, the jitter of a video frame refers to the phenomenon that the content presented by the video frame jitters or shakes during video playback. Exemplarily, when a user holds the mobile phone 100 to take a picture and moves the mobile phone 100 when hoping to shoot another scene, there may be an obvious jitter situation. The clarity of a video frame refers to the clarity of each fine texture and its boundary in the video frame. The pixels of a video frame can represent the brightness of the video frame.
[0233] It can be understood that when a video frame is compared with its adjacent video frames, the value corresponding to its jitter degree, the degree of change of other image parameters, and the degree of change of the main type are positively correlated with the degree of change of the content presented by the video. That is, the greater the change in the value corresponding to the jitter degree, the more obvious the change in the degree of change of the image parameters and the main type, and the more obvious the degree of change of the content presented by the video.
[0234] In some embodiments, the video segmentation algorithm can be represented by the following formula 1. Based on the main type of the video frame and the image parameters, the segmentation score of the video frame is calculated by combining the following formula 1, and then it is determined whether to determine it as the start frame or the end frame of a video segment according to the segmentation score of the video frame. Formula 1 is as follows:
[0235] y = α×frame A +β×frame B +γ×frame Y +δ×frame T
[0236] Among them, y represents the segmentation score of the video frame, frame A represents the jitter degree score of the video frame, frame B represents the clarity change score of the video frame, frame Y represents the label change score of the video frame, frame T represents the pixel change score of the video frame, and α, β, γ, and δ respectively represent frame A 、frame B 、frameY and the frame T In some embodiments, α, β, γ, and δ may be values preset manually according to the influence degrees of the jitter degree, clarity change, label change, and pixel change of the video frame on the segmentation score of the video frame.
[0237] It should be noted that the process of determining the segmentation score of the video frame based on the subject type and multiple image parameters above is only an example. It may also be to determine the segmentation score of the video frame based on one or more of the subject type and multiple image parameters, and the present application does not limit this.
[0238] In some embodiments, video jitter detection methods such as the optical flow method, feature point matching method based on image displacement, and based on image gray distribution characteristics can be used to detect the jitter degree of multiple video frames in the video. Then, based on the correspondence between the preset jitter degree range and the jitter degree score, the jitter degree score of the video frame is determined.
[0239] In some embodiments, a clarity detection tool can be used to determine the clarity of multiple video frames in the video, and then the clarity of the video frame is compared with the clarity of the adjacent video frame to determine the clarity change value of the video frame. Subsequently, based on the correspondence between the preset clarity change value range and the clarity change score, the clarity change score of the video frame is determined. Exemplarily, the video quality detection tool can be open-source software such as FFmpeg, Video Quality Measurement Tool, etc.
[0240] In some embodiments, the subject type of a video frame can be compared with the subject type of the adjacent video frame, and the comprehensive label change of the video frame is determined based on the quantity change of the subject type of the video frame and the content change of the subject type. Exemplarily, the union and intersection of the subject type of a video frame and the subject type of its adjacent video frame can be calculated. The quantity change of the subject type in the union can reflect the quantity change of the subject type of the video frame, and the quantity change of the subject type in the intersection can reflect the content change of the subject type. When there is a change in the quantity of the subject type in the union or the quantity of the subject type in the intersection, the label change score of the video frame can be determined according to the correspondence between the preset quantity change range of the subject type and the label change score.
[0241] In some embodiments, a pixel detection tool can be used to determine the pixels of multiple video frames in the video, and then the pixels of the video frame are compared with the pixels of the adjacent video frame to determine the pixel change value of the video frame. Subsequently, based on the correspondence between the preset pixel change value range and the pixel change score, the pixel change score of the video frame is determined. Exemplarily, the pixel detection tool can be plugins such as PixelStick, MeasureIt, Guides, etc.
[0242] In some embodiments, the degree of change in the content presented by two adjacent video frames in a video may be positively correlated with the magnitude of the segmentation score of the video frame, that is, the larger the segmentation score of the video frame, the more obvious the degree of change in the content it presents compared to the adjacent video frames.
[0243] Exemplarily, a video frame can be compared with its previous video frame, and a segmentation score threshold can be preset, and the video frame whose segmentation score exceeds the segmentation score threshold is determined as the starting frame of a video segment. As shown in the above formula 1, the higher the jitter degree of the video frame, the higher the value of frame A corresponding; the greater the change in clarity of the video frame compared to the previous video frame, the higher the value of frame B corresponding; the greater the change in labels of the video frame compared to the previous video frame, the higher the value of frame_Y corresponding; the greater the change in labels of the video frame compared to the previous video frame, the higher the value of frame Y corresponding; the greater the change in pixels of the video frame compared to the previous video frame, the higher the value of frame T corresponding.
[0244] Also exemplarily, a video frame can be compared with its next video frame, and a segmentation score threshold can be preset, and the video frame whose segmentation score exceeds the segmentation score threshold is determined as the ending frame of a video segment.
[0245] In some embodiments, the degree of change in the content presented by two adjacent video frames in a video may also be negatively correlated with the magnitude of the segmentation score of the video frame, that is, the smaller the segmentation score of the video frame, the more obvious the degree of change in the content it presents compared to the adjacent video frames, that is, the lower the segmentation score of the video frame, the greater the possibility of determining it as the starting frame or the ending frame of a video segment.
[0246] Exemplarily, a video frame can be compared with its previous video frame, and a segmentation score threshold can be preset, and the video frame whose segmentation score is lower than the segmentation score threshold is determined as the starting frame of a video segment. As shown in the above formula 1, the higher the jitter degree of the video frame, the lower the value of frame A corresponding; the greater the change in clarity of the video frame compared to the previous video frame, the lower the value of frame B corresponding; the greater the change in labels of the video frame compared to the previous video frame, the lower the value of frame Y corresponding; the greater the change in labels of the video frame compared to the previous video frame, the lower the value of frame Y corresponding; the greater the change in pixels of the video frame compared to the previous video frame, the lower the value of frame T corresponding.
[0247] Step 3: The multimodal understanding module uses computer vision services to determine the representative frame corresponding to each video segment.
[0248] After obtaining multiple video segments, the multimodal understanding module determines a representative frame from each video segment that can represent the video semantics of that video segment.
[0249] In some embodiments, the representative frame can be the starting frame, ending frame, middle frame, or random frame in a video segment.
[0250] Exemplarily, assume a video segment contains 99 video frames. Then, the 1st video frame (starting frame), the 99th video frame (ending frame), or the 50th video frame (middle frame) among these 99 video frames can be used as the representative frame of this video segment according to the time order. Alternatively, a video frame randomly selected from these 99 video frames (random frame) can be used as the representative frame of this video segment.
[0251] In some embodiments, an optimal frame can be determined from a video segment as the representative frame based on preset rules related to the image parameters and subject type of the video frame.
[0252] Exemplarily, based on the image parameters and subject type of the video frame, a score of the video frame can be calculated. The more tags a video frame has, the higher its score. The lower the jitter degree of the video frame, the higher its score. The higher the clarity of the video frame, the higher its score. The lower the pixel change of the video frame compared to the pixels of the previous video frame, the higher its score. Finally, the video frame with the highest score (i.e., the optimal frame) in a video segment can be determined as the representative frame.
[0253] S504: The multimodal understanding module returns the representative frame of the video and its related information to the gallery service module.
[0254] The related information of the representative frame includes but is not limited to the time points corresponding to the starting frame, ending frame, and the representative frame in the video segment corresponding to the representative frame, the subject type of the video segment, etc. Among them, the time points corresponding to the starting frame, ending frame, and representative frame in the video segment are beneficial for jumping to the corresponding video frame when displaying video search results for users later, improving the user experience. The subject type of the video segment is beneficial for displaying more accurate video search results when conducting video searches later.
[0255] In some embodiments, the subject type of the video segment can be used to indicate the type of the object shown in the video segment. Exemplarily, the union of the subject types corresponding to the multiple video frames included in the video segment can be obtained to get the subject type of the video segment.
[0256] In addition, the relevant information representing the frame may further include the names of the people corresponding to the video segments and the relationships between the people. In some embodiments, after determining the representative frame, the names of the people corresponding to the people in the representative frame and the relationships between the people may be identified based on the preset names of the people and the relationships between the people.
[0257] S505: The gallery service module stores the representative frame of the video and its relevant information.
[0258] After receiving the representative frame of the video and its relevant information returned by the multimodal understanding module, the gallery service module may store them. In some embodiments, the gallery service module is configured with a database, and the gallery service module may store the representative frame of the video and its relevant information in the database.
[0259] S506: The gallery service module calls the multimodal understanding module to determine the visual semantic vector corresponding to the representative frame.
[0260] In some embodiments, the gallery service module calls the computer vision service provided by the multimodal understanding module to determine the visual semantic vector corresponding to the representative frame. A video includes multiple video segments, and each video segment corresponds to a representative frame, so a video corresponds to multiple visual semantic vectors. The gallery service module calls the computer vision service provided by the multimodal understanding module to perform video semantic understanding on the video, and further determines the visual semantic vectors corresponding to the multiple representative frames in the video.
[0261] In some embodiments, the computer vision service may provide a CLIP model. The representative frame may be input into the image encoder of the CLIP model to obtain the visual semantic vector corresponding to the representative frame.
[0262] In some embodiments, the image encoder of the CLIP model may be trained with image training samples, and the image training samples may include picture training samples and video frame training samples.
[0263] S507: The multimodal understanding module returns the visual semantic vector corresponding to the representative frame to the gallery service module.
[0264] S508: The gallery service module stores the visual semantic vector corresponding to the representative frame.
[0265] After receiving the visual semantic vector corresponding to the representative frame returned by the multimodal understanding module, the gallery service module may store it. Exemplarily, the gallery service module is configured with a database, and the gallery service module may store the visual semantic vector corresponding to the representative frame in the database.
[0266] S509: The gallery service module sends the representative frame of the video and its relevant information, as well as the visual semantic vector of the representative frame, to the search module.
[0267] S510: The search module constructs an index corresponding to the video.
[0268] An index is an index corresponding to a video segment in the video. A video includes multiple video segments, so a video corresponds to multiple indexes, and different indexes correspond to different video segments in the video.
[0269] The search module can combine the visual semantic vector representing the frame, the relevant information of the frame, and the attribute tags of the video segment to construct an index corresponding to the video segment of the frame. As shown in the above example, the index of the video segment can include the visual semantic vector of the representative frame of the video segment, the time point corresponding to the start frame, the time point corresponding to the end frame, and the time point corresponding to the representative frame in the video segment, the main type of the video segment, and the attribute tags of the video segment (such as the video acquisition time, the video acquisition location, and the video storage path), etc.
[0270] It should be noted that the content included in the above index of the video segment is only an example, and the index of the video segment can include any information in the visual semantic vector representing the frame, the relevant information of the frame, and the attribute tags of the video segment. This application does not make any limitations in this regard.
[0271] Exemplarily, the search module can include an index library, and the search module can store the index corresponding to the video in the index library for subsequent search and matching based on the index library.
[0272] It can be understood that in practical applications, the number of videos stored in the mobile phone 100 may be very large, and a video corresponds to multiple indexes, indicating that the index library may contain a large number of indexes. In some embodiments, in order to improve the search efficiency, the inverted index method can be used for searching.
[0273] Exemplarily, after the search module constructs the index corresponding to the video, the visual semantic vectors included in the indexes in the index library are vector-clustered to obtain multiple clustering clusters, that is, the vector space corresponding to all indexes is divided into multiple vector regions. A vector region includes a clustering cluster, and each vector region includes multiple indexes with relatively high vector similarity. Each vector region can be replaced by a clustering center point. For example, methods such as K-means or hierarchical clustering can be used for clustering. This application does not make any limitations in this regard.
[0274] In this way, when the mobile phone 100 performs search and matching based on the text semantic vector of the text information, it can first match with the clustering center point, and then match with the visual semantic vectors in the vector region to which the determined clustering center point belongs, without having to match with all the indexes in the index library, which not only saves computing resources, but also greatly reduces the time consumed by search and matching, can avoid search delay, improve the search efficiency of the video, and further enhance the user experience.
[0275] In some embodiments, the search module may store the received representative frames and their related information, as well as the visual semantic vectors of the representative frames, for convenient lookup during index construction, thereby improving the index construction speed.
[0276] Exemplarily, the visual semantic vectors of the representative frames corresponding to a video segment and the related information of the representative frames may be stored in a document. Taking video segment 1 as an example, the visual semantic vectors of the representative frames corresponding to video segment 1 and the related information of the representative frames may be stored in sub-document 1. The information stored in sub-document 1 can be seen in Table 2 below:
[0277] Table 2
[0278] Information Name Information Content segments.media-vector [bd de d1 b4 3c 8c 9c] segments.startTime 0 segments.endTime 91666 segments.startFrame 0 segments.endFrame 2750 segments.tag-name People|Scenery|Buildings
[0279] As shown in Table 2, segments.media-vector represents the visual semantic vector of the representative frame in video segment 1. In practical applications, a vector is usually composed of an array. In the embodiments of the present application, the visual semantic vector is processed in a hexadecimal sequence to obtain a visual semantic vector in the form of [bd de d1 b4 3c 8c 9c], and the hexadecimal floating-point number can enable the mobile phone 100 to store it with less storage space, which can save storage space. segments.startTime represents the start time 0ms of video segment 1 (i.e., the time point corresponding to the start frame in video segment 1). segments.endTime represents the end time 91666ms of video segment 1 (i.e., the time point corresponding to the end frame in video segment 1). segments.startFrame and segments.endFrame represent that there are a total of 2750 video frames from the start to the end in video segment 1. segments.tag-name represents the main type of video segment 1. For example, it includes "person", "scenery", and "building".
[0280] In some embodiments, the attribute tags of video segment 1 may also be stored in the sub-document 1, and the attribute tags of video segment 1 are the attribute tags of the video stored in step S502. Taking video segment 1 as an example of a video segment in the video shot by the user, the attribute tags of video segment 1 may include the shooting time, shooting location, and storage path in the mobile phone 100, etc.
[0281] In some embodiments, the attribute tags of video segment 1 can also be stored in Document 1. It can be understood that the attribute tags of a video are fixed, that is, the attribute tags corresponding to multiple video segments in a video are the same. Then, the attribute tags of a video can also be stored in Document 1, and the attribute tags of each video segment of the video can be obtained from Document 1, which can reduce the content pressure of mobile phone 100 and storage costs.
[0282] Exemplarily, Video B includes the above-mentioned video segment 1, and also includes video segment 2 and video segment 3. The visual semantic vector of the representative frame corresponding to video segment 1 and the relevant information of the representative frame are stored in Sub-document 1. The visual semantic vector of the representative frame corresponding to video segment 2 and the relevant information of the representative frame can be stored in Sub-document 2. The visual semantic vector of the representative frame corresponding to video segment 1 and the relevant information of the representative frame can be stored in Sub-document 3. Then, the information stored in Document 1 can be seen in Table 3 below:
[0283] Table 3
[0284] Information Name Information Content file_path pictures / video shooting-time 2023.09.30 location City B segments Sub-document 1, Sub-document 2, Sub-document 3
[0285] As shown in Table 3, file_path represents the storage path of Video B in mobile phone 100. shooting-time represents the shooting time of Video B. location represents the shooting location of Video B. segments is used to indicate the sub-documents corresponding to the multiple video segments included in Video B respectively. Based on Table 2 and Table 3 above, the attribute tags of video segment 1 can be obtained from Document 1, and the relevant information of the representative frame of video segment 1, etc., can be obtained from Sub-document 1. Similarly, the attribute tags of video segment 2 and video segment 3 can also be obtained from Document 1.
[0286] It should be noted that since video semantic understanding may consume a large amount of computing resources, in order to avoid affecting the user experience such as causing lags during use, steps S503 - S510 can be executed when mobile phone 100 is in the charging and screen-off state.
[0287] The image search method based on the image library provided in this application performs search matching based on the text information input by the user and the index corresponding to the video, and obtains search results. The index corresponding to the video is the index of a video segment in the video, and the index of the video segment includes at least the visual semantic vector of the representative frame in the video segment. The visual semantic vector of the representative frame can indicate the meaning expressed by the content of the picture shown in the representative frame, and it can represent the video semantics of the video segment. Therefore, the visual semantic vector of the representative frame is fully associated with the video semantics of the video segment in the video. Performing search matching based on the visual semantic vector of the representative frame can achieve the fusion interaction between the text information and the content of the picture shown in the representative frame, thereby improving the accuracy of the video search results and enhancing the user experience.
[0288] Next, in combination with Figure 13 Continue to introduce in detail the steps included in the search phase of the image search method based on the image library provided in this application.
[0289] S511: The image library service module receives the user's input operation for text information.
[0290] The user can input text information in the search interface provided by the mobile phone 100. The text information is the text for the user to describe the characteristics of the video they need. Exemplarily, the text information may include the video acquisition time, the video acquisition location, and the content of the picture shown in the video, etc. For example, the text information can be "Scenery taken last week", "Walking by the sea in August last year", and this application does not make any limitations.
[0291] S512: The image library service module sends the text information to the search module.
[0292] S513: The search module calls the multimodal understanding module to determine the text semantic vector corresponding to the text information.
[0293] The search module calls the multimodal understanding module to perform text semantic understanding on the text information and obtains the text semantic vector corresponding to the text information.
[0294] Text semantic understanding refers to enabling the mobile phone to understand the meaning expressed by the text, which is a key technology in natural language processing (NLP) technology.
[0295] In some embodiments, the multimodal understanding module provides a CLIP model, and the text information can be input into the text encoder of the CLIP model to obtain the text semantic vector corresponding to the text information.
[0296] In some embodiments, the text encoder of the CLIP model can be trained through text training samples.
[0297] S514: The multimodal understanding module returns the text semantic vector corresponding to the text information to the search module.
[0298] S515: The search module performs vector recall in the index library based on the text semantic vector corresponding to the text information.
[0299] Vector recall refers to recalling the indexes in the index library that match the text semantic vector corresponding to the text information.
[0300] In some embodiments, the vector similarity between the text semantic vector corresponding to the text information and the visual semantic vectors respectively included in multiple indexes in the index library can be calculated to obtain the vector similarity calculation results respectively corresponding to the multiple indexes, and the N indexes with higher vector similarity among the multiple vector similarity calculation results are used as the vector recall results. Wherein, N is an integer greater than 0. Exemplarily, N can be the number of preset vector recall results, such as 5, 8, or 10, etc.
[0301] Vector similarity refers to the degree of similarity between two vectors, which can be calculated by various methods. Exemplarily, the similarity degree can be determined by calculating the cosine similarity of the two vectors, or other methods can also be used. This application does not limit this.
[0302] In some embodiments, as described above, the vector space in the index library includes multiple vector regions, and each vector region includes multiple indexes with relatively high vector similarity. Then, in the embodiments of this application, the distance between the text semantic vector corresponding to the text information and the clustering center points of multiple vector regions in the index library can be calculated to determine the clustering center point with the closest distance. Then, calculate the vector similarity between the visual semantic vectors respectively included in the multiple indexes in the vector region to which the clustering center point belongs and the text semantic vector. Sort the indexes in descending order of vector similarity to obtain the inverted index chain corresponding to the clustering center point. Exemplarily, the first N indexes in the inverted index chain can be used as the vector recall results. Another exemplarily, the indexes whose vector similarity exceeds the vector similarity threshold can be used as the vector recall results.
[0303] The following combines Figure 14 to specifically illustrate the process of vector recall.
[0304] As Figure 14 shown, the inverted index library includes multiple clustering center points such as clustering center point 1. First, calculate the distance between the text semantic vector corresponding to the text information and the clustering center points of multiple vector regions in the index library to determine that clustering center point 1 is the clustering center point with the closest distance. Then, calculate the vector similarity between the visual semantic vectors of the multiple indexes in the vector region to which center 1 belongs and the text semantic vector, sort them in descending order of vector similarity to obtain the inverted index chain 1 corresponding to clustering center point 1, and select the TopN indexes from the inverted index chain 1 corresponding to clustering center point 1 as the vector recall results.
[0305] Among them, in the inverted zipper 1, the visual semantic vector corresponding to index 1 is the closest to the clustering center point 1, and the distances between index 2 and index 3 and the clustering center point 1 gradually become farther. Therefore, the selected TopN indexes are sequentially selected backward starting from index 1. N can be any integer greater than 0, and the present application does not limit this.
[0306] S516: The search module calls the natural language understanding module to identify the entities in the text information.
[0307] The search module calls the natural language understanding module to identify the entities included in the text information.
[0308] Exemplarily, the named entity recognition technology (NER) can be used to identify the entities with specific meanings in the text information. Entities can include but are not limited to time, place, person name, organization name, proper noun. Taking the text information "the scenery photographed last week" as an example, the entities in this text information include: "last week" and "scenery".
[0309] S517: The natural language understanding module returns the entities in the text information to the search module.
[0310] S518: The search module performs entity recall in the index library based on the entities in the text information.
[0311] Entity recall refers to recalling the indexes that match the entities in the text information in the index library.
[0312] In some embodiments, the index may include the relevant information representing frames in the video segment and the attribute tags of the video segment, and entities are included therein. For example, the video acquisition time, the video acquisition location, the main type of the video segment, etc. may all include entities. Taking the text information "the scenery photographed in city B last week" as an example, there are entities "city B" as the location, "last week" as the time, and the entity "scenery" related to the content of the video display screen. Then, matching can be performed among the entities corresponding to multiple indexes, and the indexes that match the entities in the text information are obtained as the entity recall results.
[0313] S519: The search module sorts the vector recall results and the entity recall results.
[0314] In some embodiments, the intersection result or the union result of the vector recall results and the entity recall results can be sorted.
[0315] Exemplarily, sorting can be performed according to the vector similarity between the vector recall result and the text information, as well as the entity matching degree between the entity recall result and the text information. For example, the vector similarity between the text semantic vector of the text information and the visual semantic vector of the recall result (vector recall result or entity recall result), and the matching degree between the entities in the text information and the entities in the recall result can be weighted and summed to obtain the comprehensive matching degree corresponding to the recall result, and the recall results (including vector recall results and entity recall results) are sorted in descending order according to the corresponding comprehensive matching degrees.
[0316] In this way, based on the text information input by the user, on the basis of searching and matching the text semantic vector, the entities in the text information are searched and matched, and the final display result order is obtained based on the comprehensive matching degree of the search results, ensuring that the videos presented to the user are more matching results with the user's text information, and further improving the user experience.
[0317] S520: The search module returns the search results to the gallery service module.
[0318] The search module returns the sorted search results to the gallery service module.
[0319] S521: The gallery service module presents the search results to the user.
[0320] It can be understood that in practical applications, the mobile phone 100 usually stores pictures and videos in the gallery application. Therefore, when searching in the search interface provided in the gallery application, both picture search results and video search results will be displayed. That is, in addition to the indexes corresponding to the video segments, the index library also includes the indexes corresponding to the pictures.
[0321] In some embodiments, the index corresponding to the picture may include, but is not limited to, the picture semantic vector, the attribute tags of the picture, etc. Exemplarily, as described above, a video frame can be regarded as an image, and a picture is also an image. Therefore, the picture semantic vector corresponding to the picture can be generated by the image encoder of the CLIP model provided by the multi-modal understanding module and returned to the gallery service module. Exemplarily, the attribute tags of the picture can be obtained by the gallery service module first receiving and storing the user's new operation or modification operation on the attribute tags of the picture. The search module then receives the picture semantic vector and the attribute tags of the picture sent by the gallery service module and constructs the index corresponding to the picture based on this.
[0322] Exemplarily, as Figure 8 shown, the first search result displayed in the search result interface 501 includes video search results and picture search results.
[0323] In some embodiments, search results can be presented based on multiple sorted indexes. As described above, the index corresponding to a video includes a visual semantic vector representing a frame, relevant information representing the frame, and an attribute label of a video segment. The index corresponding to a picture includes a picture semantic vector and an attribute label of the picture, etc.
[0324] Exemplarily, as Figure 8 shown in the search result display interface 501, the picture search results on page 502 show the image of the picture, and the video search results show the thumbnail and time points. The information corresponding to the thumbnail and time points is the information included in the index corresponding to the video search result. Taking Figure 8 video A in the shown scenario as an example, its thumbnail is the starting frame of video segment a included in the index, its first time point is the time point "02:18" corresponding to the starting frame of video segment a included in the index, and its second time point is the total duration of the video "08:32" included in the index.
[0325] Also exemplarily, the video search results and picture search results can also show their respective corresponding attribute labels. For example, the video search results can show the video shooting time, video shooting location, etc.
[0326] Furthermore exemplarily, the video search results can also display other content in the corresponding index, such as the relationship between people, names, etc. This application does not limit this.
[0327] It should be noted that the gallery service module, search module, multimodal understanding module, and natural language understanding module can also be located in the cloud server. That is, the cloud server uses the interaction of these four modules to implement the steps included in the index construction stage. In the search stage, it can be based on the interaction between an electronic device such as the mobile phone 100 and the cloud server to implement the steps included in the search stage.
[0328] In some embodiments, the mobile phone 100 can send the text information input by the user to the gallery service module of the cloud server, so that the gallery service module of the cloud server can interact with other modules to implement the steps included in the search stage. The gallery service module of the cloud server then sends the search results to the mobile phone 100, so that the mobile phone 100 can display the search results to the user.
[0329] In some embodiments, it is assumed that the total duration of video C is 2 minutes and it is shot at 30 frames per second, that is, 30 images are shot in one second. When video C is stored in the cloud server, the multi-modal understanding module in the cloud server can perform frame splitting on video C, decompose video C into individual video frames, and 3600 video frames can be obtained; the multi-modal understanding module then performs segmentation on video C, segments it in units of 1 second, and divides it into 120 video segments, with each video segment including 30 video frames (that is, 30 images shot in one second); the multi-modal understanding module scores the 30 video frames in each video segment and takes the video frame with the highest score as the representative frame of the video segment. Exemplarily, the scoring can be based on the jitter, clarity, pixels, etc. of the video frame, and the present application does not limit this.
[0330] Subsequently, the search module of the cloud server uses the visual semantic vector of the representative frame as the index corresponding to each video segment of video C, so as to subsequently match the text information with the representative frame of each video segment of video C. If the text information matches the 50th video segment of video C successfully, the thumbnail of the video returned to the user is the representative frame of the 50th video segment, and the returned time point is 50s. For the specific implementation manners of the above embodiments, reference can be made to Figure 13 the introduction in the above embodiments and will not be elaborated herein.
[0331] It should be noted that the multi-modal understanding module of the mobile phone 100 can also segment the videos stored in the gallery in units of 1 second, and the present application does not limit this.
[0332] It can be understood that in order to implement the above functions, the electronic device provided in the embodiments of the present application includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0333] Embodiments of the present application can divide the above-mentioned electronic device into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0334] In one example, please refer to Figure 15 , which shows a possible structural schematic diagram of the electronic device involved in the above embodiment. The electronic device 1200 includes: a processing unit 1210, a storage unit 1220, and a display unit 1230.
[0335] Among them, the processing unit 1210 is used to control and manage the operations of the electronic device 1200. For example, generating recommendation information, searching for image resources according to the information input by the user, determining the display content of the search result interface, etc.
[0336] The storage unit 1220 is used to save the program code and data of the electronic device 1200. For example, saving pictures and videos, etc.
[0337] The display unit 1230 is used to display the interface of the electronic device 1200. For example, displaying desktops, various user interfaces of the gallery (such as photo display interfaces, album display interfaces, search interfaces, search result interfaces), etc.
[0338] Of course, the unit modules in the above electronic device 1200 include but are not limited to the above processing unit 1210, storage unit 1220, and display unit 1230.
[0339] Optionally, the electronic device 1200 may further include an image acquisition unit 1240. The image acquisition unit 1240 is used to acquire pictures or videos.
[0340] Optionally, the electronic device 1200 may further include a communication unit 1250. The communication unit 1250 is used to support the communication of the electronic device 1200 with other devices. For example, obtaining pictures or videos from other devices, etc.
[0341] Among them, the processing unit 1210 can be a processor or a controller. For example, it can be a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The storage unit 1220 can be a memory. The display unit 1230 can be a display screen, etc. The image acquisition unit 1240 can be a camera, etc. The communication unit 1250 can include a mobile communication unit and / or a wireless communication unit.
[0342] For example, the processing unit 1210 is a processor (such as Figure 2 the processor 110 shown), the storage unit 1220 can be a memory (such as Figure 2 the memory 120 shown), the display unit 1230 can be a display screen (such as Figure 2 the display screen 170 shown). The image acquisition unit 1240 can be a camera (such as Figure 2 the camera 160 shown). The communication unit 1250 can include a mobile communication unit (such as Figure 2 the mobile communication module 130 shown) and a wireless communication unit (such as Figure 2 the wireless communication module 140 shown). The electronic device 1200 provided by the embodiments of the present application can be Figure 2 the mobile phone 100 shown. Among them, the above-mentioned processor, memory, display screen, camera, mobile communication unit, wireless communication unit, etc. can be connected together, for example, through a bus connection.
[0343] The embodiments of the present application also provide a chip system, which includes at least one processor and at least one interface circuit. The processor and the interface circuit can be interconnected through a line. For example, the interface circuit can be used to receive signals from other devices (such as the memory of an electronic device). For another example, the interface circuit can be used to send signals to other devices (such as a processor). Exemplarily, the interface circuit can read the instructions stored in the memory and send the instructions to the processor. When the instructions are executed by the processor, the electronic device can be made to execute the various steps in the above embodiments. Of course, the chip system can also include other discrete devices, and the embodiments of the present application do not make specific limitations on this.
[0344] An embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions run on the above-mentioned electronic device, the electronic device is enabled to execute each function or step that the mobile phone executes in the above method embodiment.
[0345] An embodiment of the present application also provides a computer program product. When the computer program product runs on a computer, the computer is enabled to execute each function or step that the mobile phone executes in the above method embodiment.
[0346] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0347] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0348] The unit described as a separated component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it may be located in one place, or may be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0349] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit exists physically alone, or two or more units are integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0350] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0351] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for image search based on an image library, characterized in that, the method includes: displaying a first user interface of the image library application; the first user interface includes thumbnails of multiple image resources in the image library application, and the image resources include pictures or videos; in response to a swipe gesture of the user on the first user interface, the thumbnails of the multiple image resources move positions along the direction of the swipe gesture, and a first button is displayed on the first user interface; in response to a click operation of the user on the first button, a search interface is displayed; the search interface includes a first search box, and the first search box is used to receive text information input by the user, and the text information is used to search for image resources.
2. The method according to claim 1, characterized in that, the method further includes: after the swipe gesture of the user on the first user interface stops, a second user interface is displayed; the second user interface includes a second search box.
3. The method according to claim 2, characterized in that, the method further includes: if no next swipe gesture is detected within a first time period after detecting a swipe gesture, it is determined that the swipe gesture of the user on the first user interface stops.
4. The method according to claim 2 or 3, characterized in that, the method further includes: in response to a click operation of the user on the second search box, the search interface is displayed.
5. The method according to any one of claims 1-4, characterized in that, the displaying of the first user interface of the image library application includes: displaying a third user interface of the image library application, the third user interface includes thumbnails of multiple image resources in the image library application; in response to a swipe gesture of the user on the third user interface, a first user interface of the image library application is displayed; the image resources included in the first user interface are different from the image resources included in the third user interface.
6. The method according to claim 5, characterized in that, the third user interface includes the second search box, and the method further includes: in response to a swipe gesture of the user on the third user interface, the second search box is hidden.
7. The method according to any one of claims 1-6, characterized in that, the method further includes: in response to receiving first text information input by the user in the first search box, a first search result interface is displayed; the first search result interface includes a first thumbnail of a first video, and the first thumbnail includes a first time point; in response to receiving second text information input by the user in the first search box, a second search result interface is displayed; the second search result interface includes a second thumbnail of the first video, and the second thumbnail includes a second time point, and the second time point is later than the first time point.
8. The method according to claim 7, characterized in that, the image of the first thumbnail is the image frame corresponding to the first time point in the first video, and the image of the second thumbnail is the image frame corresponding to the second time point in the first video.
9. The method according to claim 7 or 8, characterized in that, The method further includes: In response to a user's click operation on the first thumbnail, playing the first video starting from the first time point; In response to a user's click operation on the second thumbnail, playing the first video starting from the second time point.
10. The method according to claim 7 or 8, wherein, The method further includes: In response to a user's click operation on the first thumbnail, displaying a first playback interface of the first video; the first playback interface includes a plurality of annotation points, and the annotation points are used to indicate the starting positions of video segments corresponding to the first text information in the first video.
11. The method according to any one of claims 7-10, wherein, The first text information is different from the second text information. The first video includes a first image frame and a second image frame. The first text information matches the first image frame corresponding to the first thumbnail, the second text information matches the second image frame corresponding to the second thumbnail, and the second image frame is after the first image frame.
12. The method according to claim 11, wherein, The first text information is matched with the first image frame corresponding to the first thumbnail through a CLIP model, and the second text information is matched with the second image frame corresponding to the second thumbnail through a CLIP model.
13. The method according to claim 12, wherein, The matching of the first text information with the first image frame corresponding to the first thumbnail through the CLIP model includes: Inputting the first text information into the text encoder of the CLIP model to obtain a first text semantic vector; Inputting the first image frame into the image encoder of the CLIP model to obtain a first visual semantic vector; Based on the first text semantic vector and the first visual semantic vector, matching the first text information with the first image frame; wherein, the vector similarity between the first text semantic vector and the vector of the first clustering center point of the inverted index library is greater than or equal to a first threshold, the first visual semantic vector belongs to the first clustering cluster corresponding to the first clustering center point, the vector similarity between the first text semantic vector and the first visual semantic vector is greater than or equal to a second threshold, the inverted index library includes clustering clusters corresponding to multiple clustering center points respectively, and the multiple clustering center points are determined by clustering multiple visual semantic vectors in the inverted index library.
14. The method according to claim 13, wherein, The first search result interface further includes a third thumbnail of the first video. The first text information matches the third image frame corresponding to the third thumbnail. The first image frame is in the first video segment of the first video, the third image frame is in the third video segment of the first video, the entity of the first text information matches the entity of the attribute label of the first video segment, the entity of the first text information matches the entity of the attribute label of the third video segment, and the display order of the first thumbnail is before the third thumbnail; The display order of the first thumbnail and the third thumbnail is determined through the following steps: Based on the vector similarity between the first visual semantic vector and the first text semantic vector, and the matching degree between the attribute label of the first video segment and the entity of the first text information, determine the first comprehensive matching degree between the first thumbnail and the first text information; Based on the vector similarity between the third visual semantic vector and the first text semantic vector, and the matching degree between the attribute label of the third video segment and the entity of the first text information, determine the second comprehensive matching degree between the third thumbnail and the first text information; the third visual semantic vector is obtained by inputting the third image frame into the image encoder of the CLIP model; According to the order from largest to smallest comprehensive matching degree, display the first thumbnail before the third thumbnail.
15. The method according to any one of claims 1-14, wherein, the method further includes: displaying recommended information in the first search box; the recommended information is generated according to at least one of the shooting time, shooting location, people, subject type, and event of the first picture saved in the gallery application.
16. The method according to claim 15, wherein, the method further includes: performing semantic analysis on the first picture by using a semantic analysis algorithm to obtain at least one of the people, subject type, and event of the first picture.
17. The method according to claim 16, wherein, the first picture is a picture taken within a preset time period.
18. The method according to any one of claims 1-17, wherein, the method further includes: displaying prompt information on the first user interface, where the prompt information is used to prompt the user to click the first button to trigger the search for image resources in the gallery application.
19. An electronic device, wherein, it includes a memory, a processor, and a display screen; the display screen is used to display the user interface of the gallery application; the memory is coupled to the processor, the memory is used to store computer program code, the computer program code includes computer instructions, and the processor calls the computer instructions to enable the electronic device to execute the method according to any one of claims 1-18.
20. A computer-readable storage medium, wherein, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-18.
Citation Information
Cited By
Visual media search method, electronic device and storage medium
WO2025103081A1