Visual media search method, electronic device and storage medium

By introducing input boxes and thumbnails into the gallery application, combining CLIP model and video segmentation index, the problem of difficulty in video search in the gallery is solved, fast and accurate video search is achieved, and user experience is improved.

WO2025103081A9PCT designated stage expired Publication Date: 2025-08-14HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/126082
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-23
Filing Date
2024-10-21
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

When users look up pictures and videos in the gallery, search becomes difficult as the storage volume increases, and the prior art is difficult to provide convenient and fast search methods to accurately find the required content.

Method used

By introducing input boxes and thumbnails into the gallery application interface, users can enter text to search videos, trigger thumbnails and start playing videos from a specific time point, supporting a variety of text matching methods, combining CLIP model and video segmentation index to improve search accuracy and efficiency.

Benefits of technology

It realizes the rapid and accurate search of videos that match user needs in the gallery, reduces operation steps, improves user experience, and reduces index construction costs and search delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024126082_14082025_PF_FP_ABST
    Figure CN2024126082_14082025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present application are a visual media search method, an electronic device and a storage medium. The method comprises: displaying an interface of an image application that comprises an input box, wherein the input box comprises text, and the interface comprises thumbnails of videos and thumbnails of images. In this way, matching images and videos can be displayed to a user on the basis of text that is input by the user, so that user requirements are met, thereby improving the usage experience for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Visual media search method, electronic device and storage medium

[0001] The present invention claims priority to Chinese patent application number CN202311544834.4, filed with the State Intellectual Property Office of the People's Republic of China on November 17, 2023, entitled "A video search method, electronic device and storage medium", and claims priority to Chinese patent application number CN202311580640.X, filed with the State Intellectual Property Office of the People's Republic of China on November 23, 2023, entitled "A visual media search method and electronic device", and claims priority to Chinese patent application number CN202311580872.5, filed with the State Intellectual Property Office of the People's Republic of China on November 23, 2023, entitled "A visual media search method and electronic device", and claims priority to Chinese patent application number CN202311589348.4, filed with the State Intellectual Property Office of the People's Republic of China on November 23, 2023, entitled "A gallery-based image search method and electronic device", the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present application relates to the field of terminal technology, and in particular to a visual media search method, electronic device, and storage medium. Background Art

[0003] As the photographic capabilities of electronic devices continue to improve, many users use them to capture images and videos. They also use them to download various images and videos from the internet. They also use them to receive images and videos shared by other devices. These devices store these captured images and videos locally, and as they are used, the number of images and videos stored on them continues to increase.

[0004] Generally speaking, a gallery application (hereinafter referred to as the gallery) on an electronic device provides users with access to pictures and videos stored on the electronic device. Users can search for pictures and videos stored on the electronic device through the user interface provided by the gallery. As time goes by, the number of pictures and videos in the gallery increases, making it difficult to search.

[0005] How to provide users with a convenient and fast search method to help users accurately find the required pictures or videos among a large number of pictures and videos is a problem that needs to be solved.

[0006] Summary of the Invention

[0007] The present application provides a visual media search method, electronic device, and storage medium that can help users more conveniently find required pictures, videos, etc. in a gallery, thereby improving the user experience of searching for pictures and videos in a gallery.

[0008] To achieve the above objectives, this application adopts the following technical solutions:

[0009] In a first aspect, the present application provides a visual media search method, which is applied to an electronic device, which may be a mobile phone, tablet computer, laptop computer, or other device that includes a gallery application. The visual media search method includes: the electronic device displays a first interface of the gallery application, for example, the first interface may be an album display interface, a photo display interface, a time point display interface, or a creation display interface of the gallery application. The first interface includes an input box, which supports the function of searching for pictures / videos in the gallery application. The input box of the first interface includes a first text, which is the search text entered by the user, such as "a girl doing yoga", "a boy standing near a bridge", etc. The first interface includes a first thumbnail of the first video, which corresponds to a video frame in the first video. The first thumbnail includes a first time point, which is a timestamp of the first video. When the user triggers the first thumbnail, such as a click operation, the electronic device starts playing the first video from the first time point.

[0010] The electronic device displays a second interface of the gallery application, where the second interface includes an input box, which includes second text. The second interface includes a second thumbnail of the first video, where the second thumbnail includes a second time point, which is later than the first time point. When the user triggers the second thumbnail, the electronic device starts playing the first video from the second time point. That is, after the user triggers the first thumbnail or the second thumbnail, the electronic device plays the first video, but the starting time point of the playback is different, that is, the video frame displayed to the user is different.

[0011] Based on the first text entered by the user, the electronic device can search for the first video. After the user triggers the first thumbnail, the electronic device can play the first video from the first time point, and the screen content played from the first time point meets the user's needs. Based on the second text entered by the user, the electronic device can search for the first video. When the user triggers the second thumbnail, the electronic device starts playing the first video from a second time point later than the first time point, and the screen content played meets the user's needs. That is, after the user enters text, the video search results displayed by the electronic device can meet the user's needs. In this way, the electronic device can accurately display the video corresponding to the user's needs, reduce user operations, and thus improve its user experience.

[0012] In one possible implementation, the first text and second text entered by the user in the search box are different. For example, the first text may be "boy standing by the river," and the second text may be "boy standing near a bridge." The first video matches the first text, for example, the first video includes visual content related to "boy standing by the river." The first video matches the second text, for example, the first video includes visual content related to "boy standing near a bridge." In this way, based on the search text entered by the user, the electronic device can display videos that match the search text, meeting the user's needs without requiring the user to manually search and improving their user experience.

[0013] In one possible implementation, the first video includes a first video frame and a second video frame, and the second video frame is after the first video frame, that is, when the first video is played, the first video frame will be played first, and then the second video frame. The first thumbnail corresponds to the first video frame, that is, the first thumbnail displayed is the first video frame of the first video. The second thumbnail corresponds to the second video frame, that is, the second thumbnail displayed is the second video frame of the first video. The first text and the first video are matched further including: the first text and the first video frame are matched, for example, if the first text is "boy standing by the river", the picture content of the first video frame will include a boy standing by the river. The second text and the first video are matched further including: the second text and the second video frame are matched, for example, if the second text is "boy standing near the bridge", the picture content of the second video frame will include a boy standing near the bridge.

[0014] In this way, the search text entered by the user matches the picture content displayed by the video frame corresponding to the thumbnail, and the search text reflects the user's needs, so the electronic device can search for videos that accurately correspond to the user's needs, which can improve the accuracy of video search and thus enhance the user experience.

[0015] In one possible implementation, when a user triggers the first thumbnail of a first video and the electronic device starts playing the first video from a first time point, the electronic device may start playing the first video from the first video frame or the first start frame, where the first start frame is before the first video frame. That is, the electronic device may start playing from the first video frame that matches the first text, or may start playing from the first start frame and then play to the first video frame that matches the first text.

[0016] Similarly, when the user triggers the second thumbnail of the first video, when the electronic device starts playing the first video from the second time point, the electronic device can start playing from the second video frame that matches the second text, or start playing from the second starting frame before the second video frame, and can play to the second video frame, and the second starting frame is after the first video frame.

[0017] After the user triggers the thumbnail, the electronic device can start playing from the video frame that matches the search text, or it can start playing from a previous video frame. At the same time, the second starting frame is before the second video frame and after the first video frame, and the first video frame and the second video frame match different first texts and second texts respectively. This means that if the user triggers the second thumbnail, the first video frame that matches the first text will not be shown to the user, but the video can be played from the second starting frame that is closer to the second video frame. Compared with the first video frame, the screen content displayed by the second starting frame is more closely related to the screen content displayed by the second video frame. In this way, the screen content displayed by the video played to the user can be more coherent and more accurately correspond to user needs, further enhancing the user experience.

[0018] In one possible embodiment, the visual media search method further includes: the electronic device displays a third interface of the gallery application, the third interface including an input box, the input box including third text. The third text is different from the first text. For example, the first text may be "boy standing by the river", and the third text may be "boy standing outdoors". The third interface includes a first thumbnail of the first video. When the user triggers the first thumbnail, the electronic device starts playing the first video from the first time point.

[0019] In actual applications, different users may have different text descriptions for the same video, and the same user's text description for the same video may also change, that is, the user enters different search texts, but may want to search for the same video search results. In this application, the first text can search for the first video, and the third text can search for the first video. The electronic devices all display the first thumbnail, that is, the electronic devices can search for the same video search results based on different search texts. In this way, the electronic devices can search for the same video that accurately corresponds to the user's needs based on different search texts, which can improve the accuracy of video search and enhance the user experience.

[0020] In one possible implementation, the first interface further includes a third thumbnail of the first video, and the third thumbnail includes a third time point. The visual media search method further includes: when the user triggers the third thumbnail, the electronic device starts playing the first video from the third time point. The first text matches the third video frame corresponding to the third thumbnail, and the third time point is later than the second time point. As mentioned above, the first text can also match the first video frame corresponding to the first thumbnail. That is, the same search text can search for different video frames of the same video. In this way, different video frames of a video can be searched based on the same search text, so the present application can display search results that accurately correspond to the search text, thereby improving the accuracy of video search and thus enhancing the user experience.

[0021] In one possible embodiment, when the user triggers the third thumbnail of the first video, when the electronic device starts playing the first video from a third time point, the electronic device can start playing from the third video frame that matches the first text, or start playing from the third starting frame before the third video frame, so that the electronic device can play to the third video frame, and the third starting frame is after the second video frame.

[0022] In this way, the first video can be played starting from the third starting frame that is more closely associated with the third video frame, or the first video can be played directly from the third video frame. Both can display the screen content matching the first text to the user, meet user needs, and enhance the user experience.

[0023] In one possible implementation, the first interface further includes a fourth thumbnail of the second video, the fourth thumbnail including a fourth time point. The visual media search method further includes: when the user triggers the fourth thumbnail, the electronic device starts playing the second video from the fourth time point. The first text matches the fourth video frame corresponding to the fourth thumbnail. In this way, based on the same search text entered by the user, different video frames of different videos that match the search text can be searched, thereby improving the accuracy of video search and thereby enhancing the user experience.

[0024] In one possible implementation, the visual media search method further includes: the electronic device displays a negative one screen interface, the negative one screen interface including a search box, the search box supporting online searches and searches of local files of the electronic device; the search box including first text; the negative one screen interface including a first thumbnail of a first video, the first thumbnail including a first time point; and when a user triggers the first thumbnail, the electronic device may start playing the first video from the first time point.

[0025] In this way, this application not only supports users to search for videos through the search box of the gallery application, but also supports users to search for videos through the search box of the negative one screen interface, thereby further improving the user experience.

[0026] In a second aspect, the present application provides a visual media search method, which is applied to an electronic device, such as a mobile phone, tablet computer, or laptop computer, that includes a gallery application. The visual media search method comprises: the electronic device displays a first interface of the gallery application, such as an album display interface, a photo display interface, a time point display interface, or a creation display interface. The first interface includes an input box that supports searching for images / videos in the gallery application. The input box of the first interface includes first text, which is the search text entered by the user, such as "boy standing outdoors" or "girl dancing indoors." The first interface includes a first thumbnail of a first video. The first thumbnail includes a first time point, which is a timestamp in the first video. The first text is matched to the first video frame corresponding to the first thumbnail using a CLIP model, that is, the textual semantics of the first text are matched to the visual semantics of the first video frame. The CLIP model is a pre-trained neural network model for matching images and text, and can match the first text entered by the user with the first video frame. When the user triggers the first thumbnail, such as by clicking on it, the electronic device begins playing the first video from the first time point.

[0027] The electronic device displays the second interface of the gallery application, which includes an input box, and the input box includes a second text. The second interface includes a second thumbnail of the first video, and the second thumbnail includes a second time point, and the second time point is later than the first time point. The second text and the second video frame corresponding to the second thumbnail are matched through the CLIP model, that is, the text semantics of the second text are matched with the visual semantics of the second video frame. When the user triggers the second thumbnail. The electronic device starts playing the first video from the second time point, that is, after the user triggers the first thumbnail or the second thumbnail, the electronic device plays the first video, but the starting time point of the playback is different, that is, the video frames started to be displayed to the user are different.

[0028] This allows users to search for matching video frames based on the search text they entered. The semantics of the search text are fully correlated with the visual semantics of the video frames, enabling a fusion of the search text and the visual content displayed by the video frames, thereby improving the accuracy of video searches and enhancing the user experience.

[0029] In one possible implementation, dividing a video can produce several video segments. A first video frame is in the first video segment of the first video, and a second video frame is in the second video segment of the first video. The first and second video frames can be determined by the following steps: the electronic device first performs a first processing on the first video, such as de-framing the first video to obtain multiple video frames of the first video and a classification label for each video frame. The classification label refers to the type of object depicted in the video frame. For example, the classification label can be a person, plant, animal, building, or natural scenery. The electronic device then performs a second processing on the first video based on the classification labels corresponding to the multiple video frames, such as segmenting the first video to divide the first video into multiple video segments, resulting in multiple video segments including the first video segment and the second video segment. The electronic device then determines a video frame of the first video segment as a first video frame, i.e., the first video frame is determined as a representative frame of the first video segment, which can be used to represent the first video segment. The electronic device also determines a video frame of the second video segment as a second video frame, i.e., the second video frame is determined as a representative frame of the second video segment, which can be used to represent the second video segment.

[0030] In this way, if this solution is implemented on an electronic device, there is no need to build an index for each video frame of the video, thereby avoiding increasing the delay required for search and matching. Instead, an index for the first video segment can be built based on the first video frame, and an index for the second video segment can be built based on the second video frame, thereby greatly reducing the number of indexes, increasing the speed of video search, further improving user experience, and saving the cost of building the index.

[0031] In one possible implementation, when the electronic device performs a second processing on the first video based on the classification labels corresponding to the multiple video frames, the electronic device further includes: the electronic device performs a second processing on the first video based on the difference in image parameters corresponding to the multiple video frames of the first video, and the electronic device performs a second processing on the first video based on the difference in classification labels corresponding to the multiple video frames of the first video, wherein the image parameters are display characteristics of the video frames, such as jitter, clarity, pixel value, etc. For example, the electronic device can perform a second processing on the first video based on the difference in clarity corresponding to the adjacent video frames such as the first video frame and the second video frame of the first video, and the difference between the classification labels corresponding to the adjacent video frames such as the first video frame and the second video frame. In this way, multiple video segments including the first video segment and the second video segment are obtained.

[0032] In this way, based on the classification labels, the changes in the types of objects displayed by multiple video frames can be determined, based on the image parameters, the changes in the image quality displayed by multiple video frames can be determined, and video frames with similar displayed picture content can be divided into one video segment, which is conducive to determining the video frames that can represent the video segment, so that the representative frames corresponding to multiple video segments of a video can more completely represent the picture content displayed in the video.

[0033] In one possible implementation, the electronic device determines a video frame of the first video segment as the first video frame, that is, the representative frame that can represent the first video segment, further including: the electronic device can determine the first starting frame of the first video segment, that is, the first video frame, the first ending frame, that is, the last video frame, a random video frame of the multiple video frames of the first video segment, or the video frame corresponding to the time point of the first position as the first video frame. The time point of the first position can be obtained based on the average value of the starting time point and the ending time point of the first video segment, that is, the middle time point of the first video segment is obtained, and the video frame corresponding to the middle time point is the middle frame of the first video segment. In this way, any video frame of the video segment can represent the video segment, which is convenient for the electronic device to build an index for the video segment.

[0034] In one possible implementation, the electronic device determines a video frame of the first video segment as the first video frame, further comprising: the electronic device determining the first video frame based on image parameters corresponding to the multiple video frames of the first video segment, and classification labels corresponding to the multiple video frames of the first video segment. Based on the image quality of the video frame and the type of display object of the video frame, the electronic device determines a video frame that is more representative of the first video segment. In this way, the accuracy of subsequent indexes established based on representative frames can be improved, further improving the accuracy of video searches and enhancing the user experience.

[0035] In one possible implementation, matching the first text with the first video frame corresponding to the first thumbnail using the CLIP model further includes: the electronic device inputting the first text into a text encoder of the CLIP model to obtain a text semantic vector for the first text, which is capable of representing the semantic features of the entire first text. The electronic device then inputs the first video frame into an image encoder of the CLIP model to obtain a visual semantic vector for the first video frame, which is capable of representing the semantic features of the first video frame. The electronic device then matches the first text with the first video frame based on the text semantic vector and the visual semantic vector.

[0036] In this way, the first video frame can be used to represent the first video segment, and its visual semantic vector is fully associated with the first video segment. Based on the text semantic vector of the first text, the electronic device matches the visual semantic vector of the first video frame, enabling a fusion interaction between the search text and the content displayed by the representative frame, thereby improving the accuracy of video search and enhancing the user experience.

[0037] In one possible implementation, the first vector similarity between the text semantic vector of the first text and the visual semantic vector of the first video frame is greater than or equal to a first threshold. Thus, the first threshold is set to identify the visual semantic vector of the first video frame with a high vector similarity, thereby displaying more accurate video search results to the user.

[0038] In one possible implementation, the inverted index library includes multiple visual semantic vectors, and the electronic device can cluster them to determine multiple cluster centers and cluster clusters corresponding to the multiple cluster centers. The visual semantic vector of the first video frame belongs to the first cluster cluster corresponding to the first cluster center. Therefore, the second vector similarity between the text semantic vector of the first text and the vector of the first cluster center of the inverted index library is greater than or equal to the second threshold. The third vector similarity between the text semantic vector of the first text and the visual semantic vector is greater than or equal to the third threshold.

[0039] In this way, when an electronic device searches and matches based on the text semantic vector of the search text, it can first match it with multiple cluster centers, and then match it with the visual semantic vectors of the clusters of the determined cluster centers, avoiding matching with all indexes in the index library. This further reduces search delays, improves video search efficiency, and ultimately enhances the user experience.

[0040] In one possible implementation, entities in the first text are matched with entities in the attribute tags of the first video segment. Thus, based on the search text input by the user, the electronic device searches and matches entities in the search text in addition to searching and matching the text semantic vectors. This means that the electronic device can perform both vector and entity recall, and can display video search results that are both vector and entity recall results to the user, further improving the accuracy of video searches and enhancing the user experience.

[0041] In one possible implementation, the first interface displayed by the electronic device also includes a third thumbnail of the second video, and the third video frame corresponding to the third thumbnail is in the third video segment of the second video. The entity of the first text is matched with the entity of the attribute label of the third video segment. In this way, the electronic device can perform entity recall in addition to vector recall, and can display both entity recall results and vector recall results to the user, thereby enriching the video search results, further improving the accuracy of video search, and enhancing the user experience.

[0042] In one possible implementation, the first interface displayed by the electronic device also includes a fourth thumbnail of the first video, the first text matches the fourth video frame corresponding to the fourth thumbnail, the fourth video frame is in the fourth video segment of the first video, and the entity of the first text matches the entity of the attribute tag of the fourth video segment. Furthermore, the entity of the first text matches the entity of the attribute tag of the first video segment, and the first text matches the first video frame. In the first interface displayed by the electronic device, the first thumbnail is displayed before the fourth thumbnail.

[0043] The display order of the first thumbnail and the fourth thumbnail can be determined by the following steps: the electronic device determines a first comprehensive matching degree between the first thumbnail and the first text based on the first vector similarity between the visual semantic vector of the first video frame and the text semantic vector of the first text, and the first matching degree between the attribute label of the first video segment and the entity of the first text. The electronic device then determines a second comprehensive matching degree between the second thumbnail and the first text based on the second vector similarity between the visual semantic vector of the fourth video frame and the text semantic vector of the first text, and the second matching degree between the attribute label of the fourth video segment and the entity of the first text. The electronic device then displays the first thumbnail before the fourth thumbnail according to the order of comprehensive matching degrees from large to small, that is, the first comprehensive matching degree is greater than the second comprehensive matching degree, and the thumbnail with a higher comprehensive matching degree is displayed in front.

[0044] In this way, the video search results are sorted based on the comprehensive matching degree, ensuring that the video search results ranked in a better position are the results that are more matched with the user's search text, further improving the user experience.

[0045] In a third aspect, the present application provides a visual media search method, which is executed by an electronic device, and the method includes: displaying a first interface, the first interface including a search box; receiving a first operation of a user entering a search text in the search box; the search text including first information representing a relationship between characters; displaying a first search result in response to the first operation; the first search result corresponds to a first video, the first video including a first character, the relationship between the first character and the central character is consistent with the first information, and the central character refers to a character who is at the social center among the characters contained in the visual media stored in the electronic device.

[0046] The third aspect provides a visual media search method. When a user inputs search text that includes first information representing character relationships, the electronic device can respond to the user's operation and display a first video containing a character relationship that matches the first information. This not only includes videos in the search results, but also provides users with more comprehensive search results and an improved user experience. Furthermore, users can search for matching videos based on character relationships, providing users with a more multi-dimensional search method and enhancing the user experience.

[0047] In one possible implementation, the method further includes: in response to the first operation, displaying a second search result, the second search result corresponding to the first image; the first image includes a second person, and the relationship between the second person and the central person is consistent with the first information.

[0048] In other words, search results include not only videos but also images, thus providing users with more comprehensive search results and improving user experience.

[0049] In a possible implementation, the first video includes multiple second video segments, the multiple second video segments include at least one hit video segment, and the hit video segment is a video segment containing the first person in the multiple second video segments.

[0050] In this implementation, the video is segmented to generate multiple segments. Since videos are dynamic, they often contain many scene changes. By segmenting different scenes, this approach not only enables the recognition of multiple image semantics and character relationships within the video, but also provides more comprehensive and accurate information. Furthermore, for each video segment, it reduces the presence of irrelevant information, improving search accuracy.

[0051] In one possible implementation, the method also includes: receiving a second operation of a user clicking on the first search result; in response to the second operation, playing the first video from the first frame image of the first hit video segment, or, playing the first video from the representative frame image; the first hit video segment is one of at least one hit video segment, the representative frame image is a frame image in the first hit video segment, and the representative frame image contains the first person.

[0052] In this implementation, when a user clicks on the first search result, that is, the user clicks on the first video, the first video starts playing from a certain frame of the hit video segment in response to the user's click. In this way, the first video can be played directly from a position related to the search text, and the user can directly see the hit video segment related to the search text without having to wait for playback or drag the playback progress bar to find the hit video segment, thereby improving the user experience.

[0053] In one possible implementation, the first search result is displayed, including: displaying a thumbnail of a cover frame image; the cover frame image is a frame image in the first video, and the thumbnail displays the first playback time, which is located within a closed interval formed by the start playback time and the end playback time of the first hit video segment.

[0054] In this implementation, the thumbnail of the cover frame image displays the first playback moment, so that the user can intuitively know the position of the hit video segment in the first video, thereby improving the user experience.

[0055] In a possible implementation, the first playback moment is: the start playback moment of the first hit video segment, the end playback moment of the first hit video segment, the middle playback moment of the first hit video segment, or the playback moment corresponding to the representative frame image.

[0056] In a possible implementation, the cover frame image is the first frame image of the first video, or the first frame image of the first hit video segment, or a representative frame image.

[0057] In this implementation, the cover frame is the first frame image or representative frame image of the first hit video segment. The user can intuitively know the scenes in the video related to the search text through the cover frame, thereby improving the user experience.

[0058] In a fourth aspect, the present application provides a visual media search method, which is executed by an electronic device, and the method includes: displaying a first interface, the first interface including a search box; receiving a first operation of a user entering a search text in the search box; in response to the first operation, identifying first information representing a relationship between characters in the search text; searching for the first video in multiple videos based on the first information; the first video includes multiple second video segments, the multiple second video segments include at least one hit video segment, and the character relationship information corresponding to the hit video segment is consistent with the first information; the character relationship information is used to represent the relationship between the characters contained in the visual media and the central character, and the central character refers to the character who is at the social center among the characters contained in the visual media stored by the electronic device; and displaying the first search result corresponding to the first video.

[0059] A fourth aspect provides a visual media search method. After a user enters search text, the electronic device can identify and search for information within the text, including first information representing character relationships within the text. The electronic device can then search for videos corresponding to segments containing characters whose relationships match the first information, obtaining the first video. This method enables video search, providing users with more comprehensive search results and improving the user experience. Furthermore, users can search for matching videos based on character relationships, providing a more dimensional search method and enhancing the user experience. Furthermore, this method searches based on multiple video segments. Videos are dynamic, often with many scene changes within a single video. Different video segments can represent different scenes. This method not only identifies multiple image semantics and character relationships within the video, but also provides more comprehensive and accurate information. Furthermore, for each video segment, the disturbance of irrelevant information can be reduced, improving search accuracy.

[0060] In one possible implementation, the method further includes: searching for a first image among multiple images based on the first information; character relationship information corresponding to the first image is consistent with the first information; and displaying a second search result corresponding to the first image.

[0061] In other words, search results include not only videos but also images, thus providing users with more comprehensive search results and improving user experience.

[0062] In one possible implementation, the method also includes: performing natural language understanding processing on the search text to identify the first semantic subject and category information in the search text; performing text semantic understanding on the search text to obtain a text semantic vector; searching for the first video in multiple videos based on the first information, including: performing multi-way recall in the index library based on the first information, the first semantic subject, the category information and the text semantic vector to obtain a recall result, the recall result including the first video; the index library includes attribute information of multiple video segments and multiple images, character relationship information, and an index of visual semantic vectors, and the attribute information includes at least one of the visual media acquisition time, the visual media acquisition location, the visual media classification label, and the visual media semantic subject.

[0063] This implementation matches the search text from multiple dimensions, i.e., a comprehensive search across multiple dimensions. This significantly improves the match between search results and the search text, enhancing the user experience. Furthermore, sorting the multi-channel recall results allows users to prioritize visual media that closely matches the search text, further enhancing the user experience.

[0064] In one possible implementation, the method also includes: receiving a third operation of the user to add and / or modify the target visual media; the target visual media includes a target video and a target image; in response to the third operation, segmenting the target video to obtain multiple third video segments; extracting a frame image from each third video segment as a representative frame image; respectively obtaining the character relationship information and visual semantic vector of each target image and each representative frame image; respectively obtaining the attribute information of each target image and the attribute information of each target video; based on the character relationship information and visual semantic vector of each target image and each representative frame image, as well as the attribute information of each target image and each target video, constructing an index of the attribute information, character relationship information, and visual semantic vector of each target image and each third video segment in the index library.

[0065] In this implementation, character relationship analysis and semantic understanding are performed based on representative frames extracted from each video segment. This eliminates the need to process each frame in the video segment separately, which not only reduces noise information but also reduces the amount of computation and storage, thereby reducing resource overhead.

[0066] In one possible implementation method, the character relationship information and visual semantic vector of each target image and each representative frame image are obtained respectively, including: performing character relationship analysis on each target image and each representative frame image respectively to obtain the character relationship information corresponding to each target image, and the character relationship information corresponding to each representative frame image; performing image semantic understanding on each target image and each representative frame image respectively to obtain the visual semantic vector of each target image, and the visual semantic vector of each representative frame image.

[0067] In one possible implementation, based on the character relationship information and visual semantic vectors of each target image and each representative frame image, as well as the attribute information of each target image and each target video, an index of the attribute information, character relationship information, and visual semantic vector of each target image and each third video segment is constructed in the index library, including: taking the character relationship information corresponding to the first representative frame image as the character relationship information corresponding to the fourth video segment, the first representative frame image is any representative frame image, and the fourth video segment is the video segment to which the first representative frame image belongs; taking the visual semantic vector of the first representative frame image as the visual semantic vector of the fourth video segment; based on the character relationship information and visual semantic vectors of each target image and each third video segment, as well as the attribute information of each target image and each target video, an index of the attribute information, character relationship information, and visual semantic vector of each target image and each third video segment is constructed in the index library.

[0068] In one possible implementation, image semantic understanding is performed on each target image and each representative frame image respectively to obtain a visual semantic vector of each target image and a visual semantic vector of each representative frame image, including: searching for character relationship information from the user's annotation information of the target image and the target video to obtain the annotation relationship information; identifying social circles and the types of social circles based on a face image set; the face image set includes multiple face images, a face image is an image containing a face, the multiple face images include a target face image and a face representative frame, the target face image is an image containing a face in the target image, the face representative frame is an image containing a face in the representative frame, the social circle is a set of characters having the same social relationship with the central character, and the type of the social circle represents the type of social relationship between the characters in the social circle and the central character; based on the annotation relationship information, the social circle and the type of the social circle, the character relationship information corresponding to each target image and the character relationship information corresponding to each representative frame image are determined.

[0069] In one possible implementation, social circles and their types are identified based on a collection of face images, including: performing face clustering on multiple face images to generate face clustering results; the face clustering results represent the correspondence between multiple face images, multiple faces, and multiple characters; determining a character image heterogeneous graph based on the face clustering results; the character image heterogeneous graph represents the correspondence between multiple face images and multiple characters; determining the intimacy between each two characters in a plurality of characters based on the character image heterogeneous graph, wherein the intimacy represents the degree of association between the characters; identifying social circles based on the intimacy between each two characters and the character image heterogeneous graph; and determining the type of social circle based on the character image heterogeneous graph.

[0070] In this implementation, first, social circles and their types can be identified based on face clustering results, without relying on user annotations, resulting in high intelligence and an improved user experience. Second, this process accurately identifies the central person, and then identifies the social circle and its type, without requiring information such as the user's face unlocking. This does not compromise the security of the user's private information, further enhancing the user experience. Third, the above process is based on images in a gallery database and can be completed offline without relying on a network. Compared to identifying social circles and their types based on social behavior on social networks, this method is more applicable, and the identification results are more targeted for electronic device users, further improving the user experience. Fourth, the above process first identifies social circles, then determines the social circle type based on the social circle identification results, and corrects the relationship type of the people in the social circle. Compared to simply classifying the relationships between people to obtain different social circles and social circle types, the method provided in this application has higher accuracy in identifying social circles and their types, and provides a better user experience.

[0071] In one possible implementation, the character image heterogeneous graph includes multiple image nodes and multiple character nodes, the multiple image nodes correspond one-to-one to multiple facial images, and the multiple character nodes correspond one-to-one to multiple characters; based on the intimacy between each two characters and the character image heterogeneous graph, social circles are identified, including: removing multiple image nodes in the character image heterogeneous graph, and based on the intimacy between each two characters, connecting the character nodes corresponding to two characters with non-zero intimacy through edges to obtain a character connection graph; identifying a central character based on the character image heterogeneous graph; removing the character node corresponding to the central character in the character connection graph to obtain a decentralized character connection graph; and performing community discovery on the decentralized character connection graph to obtain a social circle.

[0072] In this implementation, the heterogeneous graph of character images is simplified into a character connection graph. There is no need to find the relationship between characters based on the images in the gallery application, which simplifies the process of obtaining the character connection graph and improves the algorithm operation efficiency.

[0073] In one possible implementation, the type of social circle is determined based on a heterogeneous graph of character images, including: inputting the heterogeneous graph of character images into a graph neural network model to extract the character feature vector of each of multiple characters and the image feature vector of each facial image; inputting the character feature vector and the image feature vector into a multi-layer perceptron model to predict the type of social relationship between each of the multiple characters and the central character; determining the number of characters corresponding to each type of social relationship in the social circle; and determining the social relationship type with the largest number of corresponding characters as the type of the social circle.

[0074] In this implementation, based on the pre-trained graph neural network model and multi-layer perceptron model, the type of social relationship between the character and the central character can be predicted simply, quickly and accurately, and then the type of social circle can be determined according to the type of social relationship.

[0075] In one possible implementation, among the visual media stored in the electronic device, the time distribution divergence and location distribution divergence of the visual media containing the central character meet preset conditions, the time distribution divergence is the distribution divergence of the acquisition time of the visual media, and the location distribution divergence is the distribution divergence of the acquisition location of the visual media.

[0076] The central person is at the center of social interaction. For any person, the more evenly distributed their location and time distribution are, the more likely they are to be the central person. In this implementation, determining the central person based on whether the time distribution divergence and location distribution divergence meet pre-set conditions is more accurate.

[0077] In a fifth aspect, the present application provides a visual media search method, which is performed by an electronic device. The electronic device can be a mobile phone, tablet computer, laptop computer, or other device that includes a gallery application. The method includes: the electronic device displays a first interface of the gallery application; the first interface includes a search box; the electronic device receives a first operation in which the user enters a first text through the search box; the electronic device searches in response to the first operation to obtain multiple hit results that match the first text, the multiple hit results including at least one hit image and at least one hit video segment, and the at least one hit video segment is a video segment in at least one target video; the electronic device determines the matching score of each hit result, and the matching score represents the degree of matching between the hit result and the first text; the electronic device sorts the at least one hit image and the at least one target video according to the matching score of each hit result to obtain a first sorting result; the electronic device displays the at least one hit image and the at least one target video according to the first sorting result.

[0078] The first interface may be, for example, an album interface, a photo interface, an interface of a moment function, or a creation interface of a gallery application. The first text is a search text input by a user, such as "a girl doing yoga", "a boy standing near a bridge", and the like.

[0079] In an embodiment of the present application, the hit result includes both an image (referred to as a hit image) and a video segment (referred to as a hit video segment). The electronic device scores the matching degree of the hit image and the hit video segment respectively to obtain a matching score. Afterwards, the videos to which the hit image and the hit video segment belong (referred to as target videos, also referred to as hit videos) are sorted based on the matching score. Optionally, the higher the matching score, the higher the matching degree of the hit image or hit video segment with the first text. Optionally, the hit image and the target video can be sorted in descending order of the matching score.

[0080] The visual media search method presented on page 5 is applicable not only to image searches but also to video searches, providing users with more comprehensive search results and improving the user experience. Furthermore, this method can rank and display images and videos in the search results together, reflecting the degree of match between different images or videos and the search text. This helps users find images or videos that meet their search intent, improving the user experience.

[0081] In combination with the fifth aspect, in some implementations of the fifth aspect, at least one hit image and at least one target video are sorted according to the matching scores of each hit result to obtain a first sorting result, including: for the first target video, the highest score among the matching scores of all first hit video segments is used as the matching score of the first target video; the first target video is any one of the at least one target video, and the first hit video segment is the hit video segment in the first target video; according to the matching scores, at least one hit image and at least one target video are sorted to obtain a first sorting result.

[0082] That is to say, for any target video, the highest matching score of each hit video segment contained in the target video is used as the matching score of the target video and is included in the ranking.

[0083] A target video can include multiple matching video segments, and different matching video segments may correspond to different matching scores. In this implementation, the highest matching score among all matching video segments is used as the matching score for the target video. This better reflects the degree of matching between the target video and the first text. The target video is then ranked based on this matching score, resulting in a more accurate first ranking result, which can further improve the user experience.

[0084] In a possible implementation, there are multiple first hit video segments, and the method further includes: sorting the multiple first hit video segments according to the matching scores to obtain a second sorting result.

[0085] As described above, a target video may include multiple hit video segments. The multiple hit video segments of the same target video may be sorted, and the sorting result is referred to as a second sorting result.

[0086] In one possible implementation, after displaying at least one hit image and at least one target video according to the first sorting result, the method further includes: receiving a second operation of the user to play the first target video; in response to the second operation, playing multiple first hit video segments in sequence according to the second sorting result.

[0087] For example, a target video includes hit video segments 1, 2, and 3. These three hit video segments are sorted from highest to lowest by matching score, resulting in the following order: hit video segment 2, hit video segment 3, and hit video segment 1. When a user clicks to play the target video, the three hit video segments will play in this order. This prioritizes showing the user the hit video segments that have a high degree of match with the search text, increasing the probability of matching the user's search intent and improving the user experience.

[0088] In one possible implementation, at least one hit image and at least one target video are displayed, including: displaying thumbnails of each hit image and each target video; wherein the thumbnail of the first target video is a thumbnail of an image frame in the second hit video segment, and the second hit video segment is the video segment with the highest matching score among all first hit video segments; the thumbnail of the first target video displays one or more of the following data: the start playback time of the second hit video segment, the end playback time of the second hit video segment, the middle playback time of the second hit video segment, the total number of first hit video segments, and the total number of video segments contained in the first target video.

[0089] Optionally, the thumbnail of the first target video can be a frame from the hit video segment with the highest matching score, for example, a representative frame from the hit video segment with the highest matching score. This representative frame is highly likely to match the search text (the first text), and the user can intuitively know from the thumbnail that the target video contains the content they are searching for, thereby improving the user experience.

[0090] In addition, the above data is further displayed in the thumbnail, which makes it easier for users to intuitively obtain information about the video and the hitting video segment, further improving the user experience.

[0091] In one possible implementation, the matching score of each hit result is determined separately, including: determining multiple matching dimensions based on the first text; determining the dimension score of each hit result in each matching dimension, wherein the dimension score of the first hit result in the first matching dimension represents the degree of matching between the first hit result and the first text in the first matching dimension; the first hit result is any one of at least one hit result, and the first matching dimension is any one of the multiple matching dimensions; according to the dimension score of each hit result in each matching dimension, the weight of each matching dimension is calculated; according to the weight of each matching dimension, the dimension score of the first hit result in each matching dimension is weightedly summed to obtain the matching score of the first hit result.

[0092] Optionally, the first text may be recognized for natural language understanding, and the recognized semantic entities, character relationships, and other information categories may be used as matching dimensions. Meanwhile, the semantic vector may be used as a matching dimension.

[0093] This implementation calculates the match score based on a weighted summation approach, fully considering the contribution of information from different matching dimensions to the final match level. As a result, the resulting match score is more realistic and accurate, and the probability of different hits having the same match score is reduced, reducing the likelihood of hits being unable to distinguish the degree of match. Furthermore, this method does not require the construction of a dataset and, therefore, the collection of user data, better protecting user privacy and security, meeting the principle of "minimizing" user data, and improving the user experience.

[0094] In one possible implementation, the weight of each matching dimension is calculated based on the dimensional score of each hit result in each matching dimension, including: normalizing the dimension score matrix to obtain a normalized matrix, where the dimension score matrix is ​​a matrix composed of the dimensional scores of each hit result in each matching dimension; calculating the information entropy of each matching dimension based on the normalized matrix; calculating the weight of each matching dimension based on the information entropy of each matching dimension, wherein the weight of the first matching dimension is negatively correlated with the information entropy of the first matching dimension.

[0095] In this implementation, the information entropy of each matching dimension is calculated and weights are determined based on this entropy. Manual weighting is not required. Furthermore, information entropy can quantitatively measure the degree of data variation across different matching dimensions, reflecting the contribution of data from different matching dimensions to the final matching score. This allows accurate weighting of different matching dimensions, improving the accuracy of matching score calculations. Furthermore, by normalizing the dimension score matrix and calculating information entropy and weights based on the resulting normalized matrix, this method ensures that the matching degrees of hit results across different matching dimensions are uniformly quantified, improving the accuracy of weight and matching score calculations.

[0096] In a sixth aspect, the present application provides a visual media search method, which is executed by an electronic device, and the method includes: displaying a first interface of a gallery application; the first interface includes a search box, the search box includes a first text, and the first interface also includes a first thumbnail of a first target video; the first thumbnail displays a first playback moment, and the first playback moment corresponds to a first target frame image in the first target video; in response to a user triggering operation on the first thumbnail, playing the first target video from the first playback moment; displaying a playback progress bar; the playback progress bar displays first mark information marking the first playback moment, and second mark information marking the second playback moment, and the second playback moment corresponds to the second target frame image in the first target video; wherein, the first target frame image and the second target frame image both match the first text.

[0097] The first text is the search text entered by the user through the search box. The first target video is the video matching the first text, also known as the hit video.

[0098] The triggering operation on the first thumbnail may be, for example, clicking the first thumbnail. The first mark information and the second mark information may be video tags, or may be text, symbols, and other information.

[0099] Simply put, the first target video contains at least two frame images that match the first text, one of which is called the first target frame image, and the other is called the second target frame image. The playback time corresponding to the first target frame image is the first playback time, and the playback time corresponding to the second target frame image is the second playback time. The user clicks on the thumbnail of the first target video, and the electronic device starts playing the first target video from the first target frame image, that is, starts playing from a frame image that matches the first text. In this way, the user can see the screen that meets the search intention as quickly and directly as possible, without the user having to search one by one through the playback progress bar, thereby improving the user experience.

[0100] In addition, the first mark information and the second mark information are displayed in the playback progress bar, which helps the user to intuitively know the location of the content related to the search text in the video, thereby improving the user experience.

[0101] In combination with the sixth aspect, in some implementations of the sixth aspect, the method further includes: in response to a user operation on the second tag information, playing the first target video from the second playback moment.

[0102] For example, when a user clicks the second tag, the electronic device jumps the playback progress to the second playback moment, that is, to the second target frame image. The second target frame image matches the first text. Therefore, with this method, the user can view the image matching the search text in one step, without having to drag the playback progress bar to find related content, which facilitates user operation and further improves the user experience.

[0103] In one possible implementation, the first target video also includes at least one third target frame image, the third target frame image is different from the first target frame image and the second target frame image, and the third target frame image matches the first text; the first playback time is the earliest one among the first playback time, the second playback time and the playback times corresponding to each third target frame image; or, the first target frame image is the one with the highest degree of matching with the first text among the first target frame image, the second target frame image and at least one third target frame image.

[0104] That is to say, the first target video contains multiple target frame images that match the first text. In this case, in one implementation, the thumbnail of the first target video (i.e., the first thumbnail) can display the target frame image with the earliest playback time. When the user clicks play, the video starts playing from the target frame image with the earliest playback time. That is, the target frame images that match the first text are played in chronological order. In this way, the display of the screen conforms to the chronological order, which makes it easier for users to recall the relevant scenes of the video and improves the user experience.

[0105] As another implementation, the thumbnail of the first target video can also display the target frame image that most closely matches the first text. When the user clicks play, the video starts playing from the target frame image that most closely matches the first text. This prioritizes the display of images that best match the user's search intent, improving the user experience.

[0106] In a possible implementation, the playback progress bar further displays sorting information of the first target frame image, the second target frame image, and at least one third target frame image, where the sorting information represents the sorting of matching degrees between the frame images and the first text.

[0107] Optionally, the sorting sequence number may be displayed in the playback progress bar, for example, the numbers "1", "2", or "3". This allows the user to intuitively know the degree of match between each target frame image and the search text, thereby improving the user experience.

[0108] In a possible implementation, the sorting information is represented by one or more of color, text, graphics, and symbols.

[0109] In the seventh aspect, the present application provides a visual media search method, which is executed by an electronic device, and the method includes: displaying a first interface of a gallery application; the first interface includes a search box, the search box includes a first text, and the first interface also includes a first thumbnail of a first target video; the first thumbnail displays a first playback time; in response to a user triggering an operation on the first thumbnail, playing the first target video from the first playback time; displaying a playback progress bar; the playback progress bar displays marking information of a first playback time period and marking information of a second playback time period, the first playback time period corresponds to a first video segment in the first target video, and the first playback time period includes a first playback time; the second playback time period corresponds to a second video segment in the first target video; wherein, there is at least one frame image in the first video segment that matches the first text, and there is at least one frame image in the second video segment that matches the first text.

[0110] Simply put, the first target video contains at least two video segments that match the first text (called hit video segments), one of which is called the first video segment and the other is called the second video segment. When the user clicks on the thumbnail of the first target video, the electronic device starts playing the first target video from a certain frame in the first video segment, that is, starts playing from the video segment that matches the first text. This allows users to see the video segments that meet the search intent as quickly and directly as possible, without having to search one by one through the playback progress bar, thereby improving the user experience.

[0111] In conjunction with the seventh aspect, in some implementations of the seventh aspect, playing the first target video from the first playback moment includes: playing the first video segment from the first frame image of the first video segment; or playing the first video segment from the first target frame image in the first video segment, where the first target frame image matches the first text. Optionally, the first target frame image may be, for example, a representative frame of the first video segment.

[0112] In other words, when playing a matching video segment, playback can begin from the first frame of the matching video segment. Video segments can be divided into scenes, so a matching video segment may contain multiple frames matching the first text. This method can provide users with a more comprehensive display of images that match their search intent, improving the user experience.

[0113] Optionally, the target frame image that matches the first text in the hit video segment may be played. In this way, a screen for reviewing the search intention is intuitively presented to the user, thereby improving the user experience.

[0114] In a possible implementation, the method further includes: after the first video segment is played, playing the second video segment.

[0115] In other words, only the hit video segments can be played and the non-hit video segments can be skipped, which can save user time and improve user experience.

[0116] In one possible implementation, the playback progress bar also displays first tag information corresponding to the first video segment and second tag information corresponding to the second video segment; the method also includes: playing the second video segment in response to a user operation on the second tag information.

[0117] For example, when a user clicks the second tag, the electronic device jumps the playback progress to the second video segment. The second video segment contains at least one frame image that matches the first text. Therefore, with this method, the user can view the video segment that matches the search text in one step, eliminating the need to drag the playback progress bar to find relevant content. This facilitates user operation and further improves the user experience.

[0118] In a possible implementation, playing the second video segment includes: playing the second video segment from the first frame image of the second video segment; or playing the second video segment from a second target frame image in the second video segment, where the second target frame image matches the first text.

[0119] The beneficial effects of this implementation are similar to those of the implementation of playing the first video segment and will not be described in detail.

[0120] In a possible implementation, the playback progress bar further displays sorting information of the first video segment and the second video segment, where the sorting information represents the sorting of the matching degrees between the video segments and the first text.

[0121] In a possible implementation, the sorting information is represented by one or more of color, text, graphics, and symbols.

[0122] The beneficial effects of the above two implementation methods can be found in the method of the seventh aspect and will not be repeated here.

[0123] In an eighth aspect, a visual media-based search method is provided, comprising: displaying a first user interface of a gallery application, the first user interface including thumbnails of multiple image resources (pictures or videos) in the gallery application; in response to a user's sliding gesture on the first user interface, the thumbnails of the multiple image resources move in the direction of the sliding gesture, and a first button is also displayed on the first user interface; the user can click the first button. In response to the user clicking the first button, a search interface including a first search box is displayed, the first search box is used to receive text information entered by the user, so that image resources can be searched in the gallery application based on the text information.

[0124] In this method, the electronic device displays a first user interface of a gallery application, which is used to display pictures and / or videos saved in the gallery application. Upon receiving a sliding gesture from the user on the first user interface (the user may be manually searching for photos or videos), the electronic device displays a first button. The user can conveniently enter the search interface by clicking the first button, enter text information in the search box of the search interface, and trigger the electronic device to automatically search for image resources in the gallery. The electronic device guides the user to enter the search interface provided by the electronic device conveniently and quickly through the first button, and then the user can enter text information in the search box of the search interface to realize automatic search for image resources, thereby improving the efficiency of the user in finding pictures or videos in the gallery application.

[0125] In conjunction with the eighth aspect, in one possible implementation, after the user stops the sliding gesture on the first user interface, a second user interface is displayed, the second user interface including a second search box. In response to the user clicking the second search box, a search interface is displayed.

[0126] In this method, after the user stops manually searching for photos or videos (by swiping), the phone displays a search box, prompting the user to enter text to start an automatic search. The user taps the search box to conveniently access the search interface. Entering text in the search box triggers the phone to automatically search for images in the gallery. This provides a convenient and quick way to enter automatic search.

[0127] In a possible implementation, if a next sliding gesture is not detected within a first time period after a sliding gesture is detected, it is determined that the sliding gesture performed by the user on the first user interface has stopped.

[0128] In combination with the eighth aspect, in a possible implementation, displaying the first user interface of the gallery application includes: displaying a third user interface of the gallery application, the third user interface including thumbnails of multiple image resources in the gallery application; displaying the first user interface of the gallery application in response to a user sliding gesture on the third user interface; wherein the image resources included on the first user interface are different from the image resources included on the third user interface.

[0129] In this method, in response to the user's operation of starting the gallery application, a third user interface of the gallery application is displayed, which includes thumbnails of multiple image resources (for displaying image resources). The user can view more image resources by performing a sliding gesture on the third user interface to cause the thumbnails of the image resources to scroll in the direction of the sliding gesture. For example, in response to the user's sliding gesture on the third user interface, the electronic device displays the above-mentioned first user interface, which includes a first button. That is, when the user manually views image resources through a sliding gesture, the electronic device displays the first button to guide the user to quickly enter automatic search.

[0130] In combination with the eighth aspect, in a possible implementation, the third user interface includes the above-mentioned second search box, and the second search box is hidden in response to a sliding gesture of the user on the third user interface.

[0131] In this scenario, the third user interface includes a search box. When a sliding gesture is received from the user on the third user interface, indicating that the user needs to manually search for image resources, the electronic device hides the second search box.

[0132] In combination with the above possible implementation methods, the method provided in this application can achieve beneficial effects in the following scenarios:

[0133] The electronic device displays a third user interface, which includes a search box, and the user can enter the automatic search by clicking on the search box. If the user does not click on the search box, but performs a sliding gesture to manually search for the required pictures or videos, the electronic device hides the search box. In the process of the user manually searching for the required pictures or videos, the user may find that the manual search is too inefficient and needs to perform an automatic search. In the method provided in the embodiment of the present application, the electronic device displays a first button, guiding the user to quickly enter the automatic search by clicking the first button; it provides convenience for the user to find image resources in the gallery.

[0134] In combination with the eighth aspect, in one possible implementation, in response to receiving first text information entered by a user in a first search box, a first search result interface is displayed; the first search result interface includes a first thumbnail of the first video, and the first thumbnail includes a first time point; in response to receiving second text information entered by the user in the first search box, a second search result interface is displayed; the second search result interface includes a second thumbnail of the first video, the second thumbnail includes a second time point, and the second time point is later than the first time point.

[0135] In this method, based on the text information entered by the user, the electronic device automatically searches the gallery for videos corresponding to the text information and displays a thumbnail of the video in the search results interface. A video may contain image frames that correspond to the text information, as well as image frames that do not. In this method, the video thumbnail displays the image and time of the image frame corresponding to the text information. This allows users to easily access the video corresponding to the text information and accurately identify the specific video segment within the video that corresponds to the text information.

[0136] In combination with the eighth aspect, in a possible implementation, the image of the first thumbnail is an image frame corresponding to a first time point in the first video, and the image of the second thumbnail is an image frame corresponding to a second time point in the first video.

[0137] In combination with the eighth aspect, in one possible implementation, in response to a user clicking on a first thumbnail, the first video is played starting from a first time point; and in response to a user clicking on a second thumbnail, the first video is played starting from a second time point.

[0138] In this method, when the user triggers the first thumbnail, the electronic device starts playing the first video from the first time point. When the user triggers the second thumbnail, the electronic device starts playing the first video from the second time point. That is, after the user triggers the first thumbnail or the second thumbnail, the electronic device plays the first video, but the starting time points of the playback are different, that is, the image frames displayed to the user are different. The starting time point is the time point corresponding to the specific video segment corresponding to the text information. In other words, the electronic device directly displays the video segment related to the text information entered by the user to the user, eliminating the need for the user to manually search in the video, thereby improving the user experience.

[0139] In combination with the eighth aspect, in a possible implementation, in response to a user clicking on a first thumbnail, a first playback interface of the first video is displayed; the first playback interface includes multiple marking points, which are used to indicate the starting position of the video segment in the first video corresponding to the first text information.

[0140] In this method, each video segment corresponding to the text information entered by the user can be marked in the video playback progress bar. This allows users to easily know the starting time of all video segments corresponding to the text information entered by the user. Furthermore, users can click any marked point to have the phone start playing the video from the corresponding position of the marked point, allowing users to easily find the video segment they are looking for.

[0141] In combination with the eighth aspect, in one possible embodiment, the first text information is different from the second text information, the first video includes a first image frame and a second image frame, the first text information matches the first image frame corresponding to the first thumbnail, the second text information matches the second image frame corresponding to the second thumbnail, and the second image frame is after the first image frame.

[0142] The first text information and the first image frame corresponding to the first thumbnail are matched through the CLIP model, and the second text information and the second image frame corresponding to the second thumbnail are matched through the CLIP model.

[0143] The first text and the first image frame corresponding to the first thumbnail are matched through the CLIP model, that is, the textual semantics of the first text are matched with the visual semantics of the first image frame. The second text and the second image frame corresponding to the second thumbnail are matched through the CLIP model, that is, the textual semantics of the second text are matched with the visual semantics of the second image frame. In this way, based on the text information input by the user, image frames that match the text information can be searched. The textual semantics of the text information and the visual semantics of the image frame are fully associated, which can achieve the fusion interaction between the text information and the image content displayed by the image frame, thereby improving the accuracy of video search and enhancing the user experience.

[0144] In combination with the eighth aspect, in a possible embodiment, matching the first text information with the first image frame corresponding to the first thumbnail through the CLIP model includes: inputting the first text information into the text encoder of the CLIP model to obtain a first text semantic vector; inputting the first image frame into the image encoder of the CLIP model to obtain a first visual semantic vector; matching the first text information with the first image frame based on the first text semantic vector and the first visual semantic vector; wherein, the vector similarity between the first text semantic vector and the vector of the first cluster center point of the inverted index library is greater than or equal to a first threshold, the first visual semantic vector belongs to the first cluster cluster corresponding to the first cluster center point, the vector similarity between the first text semantic vector and the first visual semantic vector is greater than or equal to a second threshold, and the inverted index library includes multiple cluster clusters corresponding to the multiple cluster centers, and the multiple cluster centers are determined by clustering the multiple visual semantic vectors in the inverted index library.

[0145] In this way, when an electronic device searches and matches based on the text semantic vector of user-entered text information, it can first match it with multiple cluster centers, and then match it with the visual semantic vectors of the clusters of the determined cluster centers, avoiding matching with all indexes in the index library. This further reduces search delays, improves video search efficiency, and ultimately enhances the user experience.

[0146] In combination with the eighth aspect, in a possible implementation, the first search result interface also includes a third thumbnail of the first video, the first text information matches the third image frame corresponding to the third thumbnail, the first image frame is in the first video segment of the first video, the third image frame is in the third video segment of the first video, the entity of the first text information matches the entity of the attribute tag of the first video segment, the entity of the first text information matches the entity of the attribute tag of the third video segment, and the display order of the first thumbnail is before the third thumbnail.

[0147] In this method, entities in the first text information are matched with entities in the attribute tags of the first video segment, and entities in the first text information are matched with entities in the attribute tags of the third video segment. In this way, based on the text information input by the user, the electronic device not only searches and matches the text semantic vectors, but also searches and matches entities in the text information. In other words, the electronic device can perform both vector and entity recall, and can display video search results to the user that are both vector and entity recall results, further improving the accuracy of video searches and enhancing the user experience.

[0148] The display order of the first thumbnail and the third thumbnail is determined by the following steps:

[0149] Based on the vector similarity between the first visual semantic vector and the first text semantic vector, and the matching degree between the attribute label of the first video segment and the entity of the first text information, a first comprehensive matching degree between the first thumbnail and the first text information is determined; based on the vector similarity between the third visual semantic vector (obtained by inputting the third image frame into the image encoder of the CLIP model) and the first text semantic vector, and the matching degree between the attribute label of the third video segment and the entity of the first text information, a second comprehensive matching degree between the third thumbnail and the first text information is determined; and the first thumbnail is displayed before the third thumbnail in descending order of the comprehensive matching degrees.

[0150] In this way, the video search results are sorted based on the comprehensive matching degree, ensuring that the video search results ranked in a better position are the results that are more matched with the text information entered by the user, further improving the user experience.

[0151] In combination with the eighth aspect, in a possible implementation, recommended information is displayed in the first search box; the recommended information is generated based on at least one of the shooting time, shooting location, person, subject type and event of the first picture saved in the gallery application.

[0152] In this method, the electronic device displays recommended information to the user in the search box of the gallery. The recommended information is a sentence with natural semantics generated based on the first picture in the gallery. The user can enter text information to search based on the content and format of the recommended information. For example, the user can use the recommended information as text information in the search box to search. Since the recommended information is generated based on the first picture in the gallery, the electronic device can search for the first picture and all image resources with similar content to the first picture in the gallery based on the text information, thereby improving the search hit rate. Moreover, since the recommended information is a sentence with natural semantics, it can more accurately express the purpose of the user's search, narrow the scope of search results, and help users find the required pictures or videos more conveniently.

[0153] In conjunction with the eighth aspect, in a possible implementation, a semantic analysis algorithm is used to perform semantic analysis on the first image to obtain at least one of the person, subject type, and event of the first image.

[0154] In conjunction with the eighth aspect, in a possible implementation, the first picture is a picture taken within a preset time period, for example, the preset time period is more than one month from today.

[0155] In conjunction with the eighth aspect, in one possible implementation, a prompt is displayed on the first user interface, prompting the user to click the first button to trigger a search for image resources in the gallery application. In this way, the user can learn the purpose of the first button based on the prompt.

[0156] In the ninth aspect, the present application provides an electronic device comprising a memory, a display screen, and a processor; the memory stores computer program code, the computer program code comprising computer instructions; the display screen provides a display function; one or more processors call computer instructions to enable the electronic device to execute the method described in any one of the above-mentioned aspects from the first to the eighth aspect.

[0157] In a tenth aspect, the present application provides a computer-readable storage medium, in which a computer program code is stored. When the computer program code is executed by a processor, the method described in any one of the first to eighth aspects above is implemented.

[0158] In the eleventh aspect, the present application provides a computer program product, which includes: a computer program code, which, when the computer program code runs on an electronic device, enables the electronic device to execute the method described in any one of the first to eighth aspects above. BRIEF DESCRIPTION OF THE DRAWINGS

[0159] FIG1a is a schematic diagram of a search result display interface provided in an embodiment of the present application;

[0160] FIG1b is a schematic diagram of another search result display interface provided in an embodiment of the present application;

[0161] FIG1c is a schematic diagram of an application scenario of a visual media search method provided in an embodiment of the present application;

[0162] FIG1d is a schematic diagram of an application scenario of another visual media search method provided in an embodiment of the present application;

[0163] FIG1e is a schematic diagram of an application scenario of another visual media search method provided in an embodiment of the present application;

[0164] FIG2a is a diagram illustrating an example of the composition of an electronic device provided in an embodiment of the present application;

[0165] FIG2 b is a diagram illustrating an example of a software structure of an electronic device provided in an embodiment of the present application;

[0166] FIG3a is an application scenario of a visual media search method provided by an embodiment of the present application;

[0167] FIG3 b is a schematic diagram of a video playback interface provided in an embodiment of the present application;

[0168] FIG3c is a schematic diagram of a search interface provided in an embodiment of the present application;

[0169] FIG4a is an application scenario of another visual media search method provided by an embodiment of the present application;

[0170] FIG4 b is an application scenario of another visual media search method provided by an embodiment of the present application;

[0171] FIG4c is a schematic diagram showing interface changes of another visual media search method provided in an embodiment of the present application;

[0172] FIG4 d is a schematic diagram of interface changes of a visual media search method provided in an embodiment of the present application;

[0173] FIG4e is a schematic diagram of an example search result interface provided in an embodiment of the present application;

[0174] FIG4f is a schematic diagram of another search result interface provided in an embodiment of the present application;

[0175] FIG4g is a schematic diagram of another search result interface provided in an embodiment of the present application;

[0176] FIG4h is another schematic diagram of a search result interface provided in an embodiment of the present application;

[0177] FIG4i is a schematic diagram of an interface of a video playback process provided in an embodiment of the present application;

[0178] FIG4j is a schematic diagram of an interface of another example of a video playback process provided in an embodiment of the present application;

[0179] FIG4k is a schematic diagram of an interface of another example of a video playback process provided in an embodiment of the present application;

[0180] FIG41 is a schematic diagram of a search entry provided in an embodiment of the present application;

[0181] FIG4A is a schematic diagram of a scenario example of a visual media search method provided in an embodiment of the present application;

[0182] FIG4B is a schematic diagram of a scenario example of the visual media search method provided in an embodiment of the present application;

[0183] FIG4C is a schematic diagram of a scenario example of the visual media search method provided in an embodiment of the present application;

[0184] FIG4D is a schematic diagram of a scenario example of the visual media search method provided in an embodiment of the present application;

[0185] FIG4E is a schematic diagram of a scenario example of the visual media search method provided in an embodiment of the present application;

[0186] FIG4F is a schematic diagram of a scenario example of the visual media search method provided in an embodiment of the present application;

[0187] FIG4G is a schematic diagram of a scenario example of the visual media search method provided in an embodiment of the present application;

[0188] FIG4H is a schematic diagram of a scenario example of the visual media search method provided in an embodiment of the present application;

[0189] FIG4J is a schematic diagram of a scenario example of the visual media search method provided in an embodiment of the present application;

[0190] FIG4K is a schematic diagram of a scenario example of the visual media search method provided in an embodiment of the present application;

[0191] FIG5a is a schematic diagram of a visual media search method provided in an embodiment of the present application;

[0192] FIG5 b is a signaling interaction diagram of a visual media search method provided by an embodiment of the present application;

[0193] FIG6 is a schematic diagram of a process for determining a representative frame according to an embodiment of the present application;

[0194] FIG7 is a schematic diagram of a vector recall process provided in an embodiment of the present application;

[0195] FIG8 is a schematic diagram showing the principle of a visual media search method provided in an embodiment of the present application;

[0196] FIG9 is a flowchart of an example of a visual media search method provided in an embodiment of the present application;

[0197] FIG10 is a flow chart of another example of a visual media search method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0198] The following describes the technical advantages of the visual media search method provided by this application, in conjunction with related technologies. For ease of understanding, this is illustrated using an example scenario. In this example scenario, the electronic device is a mobile phone, and the mobile phone's gallery application stores multiple images and videos.

[0199] First, the terms used in the embodiments of the present application are explained. It should be understood that this explanation is for a clearer understanding of the embodiments of the present application and does not necessarily constitute a limitation on the embodiments of the present application.

[0200] Video frame: also known as image frame, refers to any frame of a video. A frame is a still picture in a video, and continuous frames can form a video.

[0201] Video segmentation refers to the process of dividing a video into segments. In some embodiments, the video can be first deframed to obtain individual video frames, and then a video segmentation algorithm can be used to segment the video comprising multiple video frames to obtain video segments. Specific implementation methods can be found in the following embodiments.

[0202] Representative frame: A video frame in a video segment that can be used to represent the video segment. For example, the representative frame can be the start frame, end frame, a random video frame, or the best frame with the highest score.

[0203] CLIP model: The CLIP (Contrastive Language-Image Pre-Training) model is a pre-trained neural network model for matching images and text. In some embodiments, the CLIP model's text encoder and image encoder perform contrastive learning to train a text encoder that outputs text semantic vectors for text and an image encoder that outputs visual semantic vectors for images or video frames.

[0204] Text semantic vector: A text can be input into a text encoder to obtain a vector that can represent the semantic features of the entire text. For example, the text encoder can use a model such as the Transformer commonly used in Natural Language Processing (NLP), which is not limited in this application. In an embodiment of the present application, the search text entered by the user can be input into the text encoder to obtain the text semantic vector of the search text.

[0205] Visual semantic vector: A visual semantic vector can be obtained by inputting a video frame of an image or video into an image encoder. For example, the image encoder can use a CNN model or a VIT model, which is not limited in this application. In an embodiment of this application, a representative frame of a video can be input into the image encoder to obtain a visual semantic vector for the representative frame.

[0206] Vector similarity: used to describe the degree of similarity between two vectors (e.g., between a text semantic vector and a visual semantic vector). In an embodiment of the present application, the video frames that match the search text can be determined by comparing the similarity between the text semantic vector and the visual semantic vector of the search text. For example, vector similarity can be calculated using the cosine similarity formula, but can also be calculated using other methods.

[0207] Entity: A word with a specific meaning in a text. For example, entities may include, but are not limited to, time, place, person, organization, and proper nouns in a text. In some embodiments, named entity recognition (NER) technology can be used to identify entities with specific meanings in a text, but this application does not limit this.

[0208] Image parameters: used to indicate the display characteristics of an image or video frame. For example, image parameters may include jitter, clarity, pixel value, etc. of a video frame, which is not limited in this application.

[0209] Attribute tag: information used to indicate the attributes of a video or video segment. In the embodiment of the present application, the attribute tag of a video may include the video acquisition location, video acquisition time, name, file name, or classification tag.

[0210] Visual content-related and visual content-independent: Visual content refers to the objects presented by visual media and the relationships between them. Computer vision can enable computers to possess capabilities similar to human vision, including the perception, understanding, analysis, and interpretation of visual content. Currently, the Generative Pre-trained Transformer 4 (GPT-4) model can support inputting an image into the model and outputting human-language natural language describing the important information in the image.

[0211] In the context of visual media search in this solution, information in a set of texts that is used to describe the visual content in the visual media is referred to as "visual content-related" information. Information related to visual content requires an electronic device to obtain through a text semantic understanding model. Information in a set of texts that is not related to the visual content of the visual media is referred to as "visual content-independent" information. Information irrelevant to visual content can be obtained through user annotations, pre-recording by an electronic device, and analysis and identification by a preset algorithm. In an embodiment of the present application, information irrelevant to visual content may include location, time, name, file attributes, names, tags, character relationship information, etc. Among the above-listed information irrelevant to visual content, the acquisition time, name, and file attributes can be automatically recorded by the electronic device when acquiring the visual media. Names and tags can be annotated by the user. Character relationship information can be identified through a character relationship analysis algorithm.

[0212] For example, in “the sky photographed in Beijing this year”, “this year” (time) and “Beijing” (place) are not related to the visual content of the visual media, so “this year” and “Beijing” are information irrelevant to the visual content; the object corresponding to “the sky” is perceived by human vision, and is information used to describe the visual content in the visual media, and is therefore information related to visual semantics.

[0213] In related art, when an electronic device's gallery stores a video, it also stores the video's shooting location or time, as well as keywords in the video's file name, as attribute tags, and uses the attribute tags as the video's index. Subsequently, users can enter simple search text such as time, location, or keywords in the file name in the gallery's search interface, and a search will be performed between the user-entered search text and the video's index (attribute tag), enabling video search.

[0214] Assume that the electronic device is a mobile phone 300, and its gallery stores Video 1, which has been pre-assigned the attribute tags "stars" and "this year." For example, "stars" can be the file name manually configured by the user for Video 1. "This year" is the time when Video 1 was captured. As shown in Figure 1a, when the user enters the keyword "stars" in the search box of the gallery provided by mobile phone 300, mobile phone 300 searches for videos with the attribute tag "stars" and displays the search results, including Video 1.

[0215] However, as shown in FIG1b , if the user enters a complex search text “a boy standing next to a tree” in the search box of the gallery, the mobile phone 300 cannot search for the video 1 and the user still needs to search manually.

[0216] In actual applications, users may enter complex search text such as "boy standing next to a tree" to describe the video they want based on the content of the video. However, electronic devices in related technologies only support search matching of keywords such as attribute tags, which may easily lead to the electronic device being unable to display the video required by the user. The user still needs to manually slide the scroll bar of the gallery to search, which is cumbersome and affects the user experience of mobile phone 300.

[0217] To address the aforementioned issues, embodiments of the present application provide a visual media search method that, based on user-entered search text, can search for video frames that match the search text. The textual semantics of the search text are fully correlated with the visual semantics of the video frames, enabling a fusion interaction between the search text and the content displayed by the video frames, thereby improving the accuracy of video searches and enhancing the user experience.

[0218] Exemplarily, Figure 1c is a schematic diagram of an application scenario of a visual media search method provided in an embodiment of the present application. Taking the electronic device as a mobile phone as an example, as shown in Figure (a) in Figure 1c, the main interface of the mobile phone includes an icon 101 of the gallery APP. The user clicks on the icon 101. In response to the user's click, the mobile phone opens the gallery APP and displays the album interface shown in Figure 1c (b). The album interface 102 includes a search box 103. The user can enter the search content in the search box 103. For example, the user can enter "Beijing" in the search box 103. In response to the user's operation, the mobile phone can search for images with the shooting location of "Beijing" in the gallery and display the search results, as shown in Figure 1c (c).

[0219] In the related art, the search function can only be applied to the search for images, not for the search for videos. Specifically, as shown in Figure 1c (c), the search results do not include videos. Moreover, the search in the related art is mainly based on the shooting location, shooting time, the name of the person or other tags that the user has annotated on the image, etc. The content that is not annotated with a tag cannot be searched, so the search results are not accurate and the user experience is not good. In addition, in the related art, it is not possible to search for images or videos based on the relationship between people, and the search experience brought to users needs to be improved. For example, as shown in Figure 1d (a), the user enters "dinner with family" in the search box 103 of the album interface 102. Even if the gallery contains images that match the sentence, the mobile phone cannot search for the results, as shown in Figure 1d (b).

[0220] Moreover, in the related art, searches are mainly based on the shooting location, shooting time, and attribute tags annotated by the user for the image. Content that is not annotated with attribute tags cannot be searched, so the search results are not accurate and the user experience is poor. For example, as shown in Figure 1e (a), when a user enters "boys standing outdoors" in the search box 103 of the album interface 102, this complex search text, although the gallery includes images or videos that match the search statement (as shown in Figure 1c (c)), the mobile phone cannot search for results, as shown in Figure 1e (b). The user still needs to manually operate the scroll bar of the sliding gallery to search, which is cumbersome and affects the user experience.

[0221] In addition, in the related art, when searching for images based on shooting time, shooting moment, etc., the search results do not involve sorting issues. If the user needs to find images that meet the actual search intention from these search results, they need to open the search results and search one by one, which is cumbersome and brings a poor search experience to the user. For example, if a user wants to search for images related to "boys standing outdoors", they cannot find them through the process shown in Figure 1e, so they can only search by shooting time or shooting moment. Assuming that the image related to "boys standing outdoors" was taken in August 2022, the user can continue to enter "August 2022" in the search box 103 in Figure (b) of Figure 1e. In response to the user's operation, the mobile phone displays the search results, as shown in Figure (c) of Figure 1e. The search results include 58 images. The user needs to click the "More" option 201 to enter the interface shown in Figure (d) of Figure 1e and search for images related to "boys standing outdoors".

[0222] To this end, an embodiment of the present application provides a visual media search method. On the one hand, the method is applicable not only to images but also to videos, providing users with more and more comprehensive search results and improving the user experience. On the other hand, the method is based on semantic vectors for search and matching, so users can perform fuzzy searches on sentences. For example, users can enter search sentences such as "woman doing yoga", and electronic devices can search based on the sentences, which can more accurately search for visual media that meets the user's actual intentions, improve the search precision, and improve the user experience. Thirdly, the method can identify character relationship information in visual media and establish a corresponding relationship between visual media and character relationships. In this way, when a user searches for a certain character relationship, images and / or videos that have this character relationship with the user (e.g., the owner of the device) can be accurately displayed to the user, providing users with more dimensional search methods and improving the user experience. Fourthly, the method can sort and display the images and videos in the search results together, reflecting the degree of match between different images or videos and the search sentence, which is more conducive to users finding images or videos that meet the search intentions and improving the user experience.

[0223] The various visual media search methods provided above can be applied to electronic devices. To facilitate understanding, the composition of the electronic device and its software structure are introduced below.

[0224] This application does not limit the type of electronic device. For example, the electronic device may be a mobile phone, tablet computer, desktop computer, laptop computer, notebook computer, ultra-mobile personal computer (UMPC), handheld computer, netbook computer, personal digital assistant (PDA), wearable electronic device, smart watch, etc. This application does not impose any special restrictions on the specific form of the above electronic devices.

[0225] In this embodiment, as shown in Figure 2a, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0226] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0227] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0228] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0229] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly retrieve it from the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.

[0230] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0231] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0232] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes the instructions stored in the internal memory 121 to perform various functional applications and data processing of the electronic device.

[0233] The internal memory 121 may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as a sound playback function during video playback, an image playback function during video playback, etc.). The data storage area may store data generated during the use of the electronic device (such as video data, etc.).

[0234] In some embodiments, the internal memory 121 stores instructions for executing a visual media search method. The processor 110 can search for videos by executing the instructions stored in the internal memory 121 .

[0235] In some embodiments, the electronic device searches for videos stored in a gallery application, where the gallery application stores videos that a user has captured using the electronic device. The electronic device can implement the capture function using an ISP, a camera 193, a video codec, a GPU, a display 194, and an application processor.

[0236] The ISP processes data fed back by camera 193. In some embodiments, the ISP can be incorporated into camera 193. Camera 193 is used to capture still images or video. A video codec compresses or decompresses digital video. This allows electronic devices to play or record video in a variety of encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.

[0237] The NPU can be used to implement intelligent cognition applications in electronic devices, such as image recognition, face recognition, voice recognition, text understanding, etc. In some embodiments, the NPU can be used to understand the search text entered by the user in the search interface provided by the electronic device.

[0238] The electronic device realizes the display function through the GPU, the display screen 194, and the application processor. The GPU is a microprocessor for image processing, which connects the display screen 194 and the application processor.

[0239] The electronic device's display screen 194 can display a series of graphical user interfaces (GUIs), which serve as the main screen of the electronic device. Generally, the electronic device's display screen 194 includes a limited number of controls, with which a user can interact through direct manipulation to read or edit information related to an application.

[0240] In some embodiments, the electronic device may include a gallery application, and the display screen 194 of the electronic device may display an icon corresponding to the gallery application. Upon user triggering, the display screen 194 of the electronic device may display a search interface for the gallery application, which may include a search control. The user may edit the search control to enter search text and trigger the control, thereby enabling the electronic device to search the gallery application based on the user-entered search text and display the search results to the user via the display screen 194.

[0241] In some embodiments, if the user triggers a video stored in the gallery application, the electronic device can implement the audio function during video playback through the audio module 170, the speaker 170A, the headphone jack 170D, and the application processor.

[0242] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The electronic device can listen to the audio of the video playing through the speaker 170A. The headphone jack 170D is used to connect a wired headset. The electronic device can listen to the audio of the video playing through the wired headset connected to the headphone jack 170D.

[0243] The charging management module 140 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also provide power to the electronic device via the power management module 141.

[0244] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.

[0245] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.

[0246] The mobile communication module 150 can provide solutions for wireless communications, including 2G / 3G / 4G / 5G, applied to the electronic device 100. In some embodiments, at least some functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 can be set in the same device as at least some modules of the processor 110.

[0247] The wireless communication module 160 can provide wireless communication solutions for application on the electronic device 100, including wireless local area networks (WLAN) (such as Wi-Fi networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc.

[0248] In the embodiment of the present application, the electronic device 100 can receive pictures or videos from other devices through the mobile communication module 150 or the wireless communication module 160, and save the received pictures or videos to the gallery.

[0249] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0250] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini-LED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc.

[0251] In the embodiment of the present application, the display screen 194 can be used to display the user interface of the electronic device 100, such as a desktop interface, an album display interface of a gallery, etc.

[0252] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.

[0253] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.

[0254] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device may include 1 or N cameras 193, where N is a positive integer greater than 1. In the embodiment of the present application, the still images or videos captured by the camera 193 can be saved to the gallery.

[0255] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when an electronic device selects a frequency, the DSP performs a Fourier transform on the frequency energy.

[0256] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. This allows electronic device 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.

[0257] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU enables intelligent cognitive applications in electronic devices, such as image recognition, face recognition, speech recognition, and text comprehension.

[0258] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.

[0259] The pressure sensor is used to sense pressure signals and can convert pressure signals into electrical signals. In some embodiments, the pressure sensor can be set on the display screen 194. There are many types of pressure sensors, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. A capacitive pressure sensor can be a device including at least two parallel plates with conductive material. When a force acts on the pressure sensor, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure based on the change in capacitance. When a touch operation is applied to the display screen 194, the electronic device 100 detects the intensity of the touch operation based on the pressure sensor. The electronic device 100 can also calculate the position of the touch based on the detection signal of the pressure sensor.

[0260] A touch sensor, also known as a "touch panel," can be provided on the display screen 194. The touch sensor and the display screen 194 form a touch screen, also known as a "touch screen." The touch sensor is used to detect touch operations applied to or near the touch sensor. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In other embodiments, the touch sensor can also be provided on the surface of the electronic device 100, in a location different from that of the display screen 194.

[0261] The electronic device 100 can detect a user's operation on the display screen 194 using a pressure sensor or a touch sensor. For example, the operation can be inputting text information in a search box, swiping up or swiping down, or clicking a control on a user interface.

[0262] In addition to the aforementioned components, electronic devices also run operating systems, such as the iOS operating system developed by Apple, the open-source Android operating system developed by Google, and the Windows operating system developed by Microsoft. Applications can be installed and run on these operating systems.

[0263] The operating system of the electronic device can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. The embodiment of the present application takes the Android system of the layered architecture as an example to illustrate the software structure of the electronic device.

[0264] FIG2 b is a block diagram of the software structure of the electronic device according to an embodiment of the present application.

[0265] A layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0266] The application layer can include a series of application packages. As shown in Figure 2b, the application package can include a gallery app, a video playback application, a search module, a natural language understanding module, and a computer vision (CV) service.

[0267] In addition, the search module, natural language understanding module, and CV service, etc., may all be located at the same layer of the electronic device, or may be located at different layers of the electronic device, or may be located at multiple layers of the electronic device at the same time, so as to realize their functions through software interfaces between layers.

[0268] The visual media search method provided in the embodiments of the present application can be implemented based on the interaction between the search module, the natural language understanding module and the CV service. The specific implementation method can be found in the detailed introduction in the embodiments below.

[0269] The gallery service module stores pictures / videos obtained through user operations such as shooting, downloading, screenshots or recording; it is also used to store representative frames of video segments, visual semantic vectors of representative frames and other information; it also receives search text entered by the user so that the electronic device can search and match videos based on the search text entered by the user; it also displays pictures / videos obtained through user operations such as shooting, downloading, screenshots or recording.

[0270] A gallery, also known as a gallery app, gallery service, or gallery service module, provides users with services such as storage, management, display, and recommendation of visual media such as images and videos. Optionally, a gallery may include a database, hereinafter referred to as the gallery database. This database stores gallery-related data, such as images and their attribute information, videos and their attribute information, visual semantic vectors, face clustering results, and person relationship information.

[0271] Optionally, the gallery APP can have multiple functions, such as album function, one-click blockbuster function, etc. The data in the gallery database can be applied to various functions of the gallery APP, and the data returned by the search module to the gallery APP can also be applied to any function of the gallery APP. This application does not impose any restrictions on this.

[0272] The video player application can be a native video player application of the electronic device. In some embodiments, the video player application can also be a third-party video player application.

[0273] The search module is used to build an index corresponding to the video based on the information stored in the above-mentioned gallery service module, and is also used to recall vectors in multiple indexes based on the text semantic vector corresponding to the search text, and to recall entities in multiple indexes based on the entities in the search text. It is also used to sort the vector recall results and entity recall results to obtain search results.

[0274] The search module is used to search for visual media based on user input. Optionally, the search module may have a database for storing index data. The index data may be indexed using an inverted index. The database of the search module is hereinafter referred to as an inverted index library.

[0275] The natural language understanding module is used to identify entities in the search text entered by the user.

[0276] The natural language understanding module is used to understand the information contained in the text. In the embodiments of the present application, the natural language understanding module can be used to identify semantic entities, character relationships, category information, etc. in the text. Category information refers to information that characterizes the type of image or video. For example, category information can include people, plants, animals, buildings, or natural scenery.

[0277] In one embodiment, as shown in FIG2 b , the CV service may include a face analysis module, a character relationship analysis module, a video segmentation module, and a multimodal understanding module.

[0278] The face analysis module is used to perform face analysis on visual media. Face analysis includes but is not limited to face recognition, face clustering, gender recognition, and age recognition.

[0279] The character relationship analysis module is used to analyze the social relationships between characters contained in visual media and obtain character relationship information.

[0280] The video segmentation module is used to segment the video, dividing it into multiple segments based on semantic content, and extracting a frame from each segment as a representative frame (also called a representative frame image). The representative frame is used to represent the semantics of the video segment.

[0281] The multimodal understanding module is used to deframe the video to obtain individual video frames, then use the video segmentation algorithm to segment the video to obtain video segments, and then determine the representative frames of the video segments and the visual semantic vectors of the representative frames.

[0282] The multimodal understanding module, also known as a multimodal semantic understanding module or a multimodal semantic understanding model, is used to perform semantic understanding on one or more of multiple modal information, such as text, images, videos, and audio. In the embodiments of the present application, the multimodal understanding module can be used to perform semantic understanding on images to obtain visual semantic vectors, and can also be used to perform semantic understanding on text (e.g., search text) to obtain text semantic vectors.

[0283] Optionally, the CV service may have its own database, hereinafter referred to as the CV database. The CV database is used to store data required for the operation of various modules in the CV service, such as images containing people, face clustering results, person relationship analysis results, etc.

[0284] Of course, in addition to the various modules or APPs shown in Figure 2b, the application layer can also include camera, call, map, Bluetooth, short message and other applications (not shown in the figure), and the embodiment of the present application does not impose any restrictions on this.

[0285] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions. As shown in Figure 2b, the application framework layer may include a window manager, a telephony manager, a content provider, a resource manager, a notification manager, a view system, and so on.

[0286] The window manager manages window programs. Content providers store and retrieve data and make it accessible to applications. This data can include videos, images, and more. The resource manager provides applications with various resources, such as localized strings, icons, images, layout files, and video files. The view system includes visual controls, such as those that display text and images. The view system is used to build applications.

[0287] In some embodiments, each application package of the application layer can implement the relevant functions of the gallery APP by calling the algorithm or module of the application framework layer. For example, when the face analysis module performs face clustering, it can cluster the faces in the image by calling the face clustering algorithm module of the application framework layer (not shown in the figure) to obtain the face clustering result. For another example, when the character relationship analysis module performs character relationship analysis, it can call the social circle recognition algorithm module of the application framework layer (not shown in the figure), identify the central person according to the face clustering result, and find the social circle of the central person. The detailed implementation of the above process will be further elaborated in subsequent embodiments.

[0288] The Android Runtime consists of a core library and a virtual machine. The Android runtime is responsible for scheduling and managing the Android system. The core library consists of two parts: one for the Java language's callable functions and the other for the Android core library.

[0289] The system library may include functional modules such as a surface manager, a media library, a 3D graphics processing library (e.g., OpenGL ES), a 2D graphics engine (e.g., SGL), and an image processing library.

[0290] The surface manager manages the display subsystem and provides fusion of 2D and 3D layers for multiple applications. The media library supports playback and recording of various common audio and video formats, as well as still image files. The 3D graphics library implements 3D graphics drawing, image rendering, compositing, and layer processing. The 2D graphics engine is the drawing engine for 2D drawing.

[0291] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.

[0292] It should be noted that although the embodiments of the present application are described using the Android system as an example, its basic principles are also applicable to electronic devices based on operating systems such as iOS and Windows.

[0293] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the application scenario of the technical solution of the present application is first described below.

[0294] First, let's introduce the application scenario 1 of video segmentation.

[0295] Exemplarily, the visual media search method provided in the embodiments of the present application can be implemented in a gallery search scenario of an electronic device.

[0296] In conjunction with Figure 3a, the following uses a mobile phone 300 as an example electronic device to exemplify the visual media search method provided by an embodiment of the present application. In this scenario, the gallery of the mobile phone 300 may include multiple pictures and multiple videos. In some embodiments, the pictures / videos may be taken by the user through the mobile phone 300 and stored in the gallery, or downloaded and stored in the gallery by the user through other application platforms, or downloaded and stored in the gallery by the user after sending other electronic devices to the mobile phone 300.

[0297] In this visual media search method, the mobile phone 300 can construct an index for the videos stored in the gallery when it is charging or in the screen-off state. In the process of constructing the index, the mobile phone 300 can deframe the video to obtain individual frames, use a video segmentation algorithm to segment the video based on multiple video frames to obtain video segments, and then determine the representative frames of the video segments and the visual semantic vectors of the representative frames, and then construct an index for the video segments based on the visual semantic vectors of the representative frames. The specific implementation method can be found in Figure 5b and the detailed description in the embodiments below. It should be understood that the pictures stored in the gallery do not need to be deframed, and the index of the pictures can be constructed by referring to the method of constructing an index based on the representative frames of the video segments.

[0298] Assume that a user takes a plurality of videos on a weekend trip through a mobile phone 300 and stores them in the mobile phone 300. The user wishes to edit the videos taken during the trip in his spare time. At this time, the user can search through the search function of the gallery of the mobile phone 300 to make the mobile phone 300 display the videos taken during the trip.

[0299] As shown in Figure 3a, the mobile phone 300 displays a main interface 310, which includes application icons for multiple applications, such as a gallery application icon 311. The user triggers the gallery application icon 311, and the mobile phone 300 launches the gallery in response to the user's triggering operation and displays the gallery's album display interface 320. The album display interface 320 includes a search box 321 and multiple albums. For example, the multiple albums may include the "All Photos" album shown in Figure 3a, the "Camera" album, the "My Favorites" album, the "Screenshots and Recordings" album, the "My Favorites" album, the "Self-Created" album, and the "Video Editing" album, etc.

[0300] In one example, a user can trigger the search box 321 included in the album display interface 320 and enter the search text "boy standing outdoors" in the search box 321. The mobile phone 300 converts the user input "boy standing outdoors" into a text semantic vector; then searches and matches based on the text semantic vector and the video index and the image index. The video index includes visual semantic vectors representing frames, and the image index includes visual semantic vectors of images. The vector similarity between the text semantic vector and the visual semantic vector can be calculated, and the picture / video corresponding to the visual semantic vector whose vector similarity exceeds the vector similarity threshold is used as the first search result. The mobile phone 300 can then display the first search result display interface 330 of the gallery. The first search result display interface 330 can display the first search result corresponding to "boy standing outdoors", which can include pictures and videos.

[0301] The index of the video in the first search result includes the visual semantic vector of the representative frame. As mentioned above, the representative frame can be used to represent the corresponding video segment. Therefore, the visual semantics of the representative frame can represent the visual semantics of multiple video frames of the video segment. In other words, the visual semantic vector of the representative frame is fully associated with the video segment. When the user enters a complex search text based on the screen content displayed by the video, the present application searches and matches the text semantic vector corresponding to the search text with the visual semantic vector of the representative frame, which can realize the fusion interaction of the search text and the screen content displayed by the video segment, improve the accuracy of the video search, and thus enhance the user experience.

[0302] In some embodiments, the first search results displayed by the mobile phone 300 may be sorted in descending order according to the degree of relevance to the search text input by the user.

[0303] As shown in Figure 3a, the first search result display interface 330 displayed by the mobile phone 300 may include a first display area 331, a second display area 332, and a third display area 333. In the first display area 331, all first search results, including pictures and videos, are displayed. In the second display area 333, a best-matching video search result is displayed, and the number of all video search results, "4", is displayed. The time point corresponding to the best-matching video search result is also magnified and displayed. In the third display area 332, a best-matching image search result is displayed, and the number of all image search results, "6", is displayed.

[0304] As shown in Figure 3a, the video search results of the first search result display interface 330 display thumbnails and time points. In some embodiments, the thumbnail corresponds to a video frame of a video segment in the video, and the visual semantic vector representing the frame in the video segment matches the text semantic vector corresponding to the search text entered by the user, that is, the vector similarity between the two exceeds the vector similarity threshold. The time point of the video can be the time point of the video segment to which the video frame corresponding to the thumbnail belongs. Exemplarily, the time point can be the time point of the video frame corresponding to the thumbnail, or it can be the time point of other video frames in the video segment.

[0305] Taking the video segment a of a video search result "Video A" in the first display area 331 as an example, "Video A" is divided into multiple video segments. The thumbnail displayed of "Video A" corresponds to the starting frame of video segment a. The vector similarity between the visual semantic vector of the representative frame of video segment a and the text semantic vector of "boy standing outdoors" exceeds the vector similarity threshold.

[0306] The second display area 332 includes two time points. The first time point is "02:18," indicating that the start frame of video segment a begins displaying at "02:18." The second time point indicates that the total duration of "Video A" is "08:32." As shown in FIG3a , if "Video A" is the video desired by the user, the user can trigger "Video A." In response to the user's triggering operation on "Video A," the mobile phone 300 displays a video playback interface 340. In this interface, the mobile phone 300 adjusts the time progress of "Video A" to "02:18" and starts playback, i.e., starting from the start frame of video segment a in "Video A."

[0307] It should be noted that the thumbnails and time points shown in the above "Video A" are only examples.

[0308] For example, the thumbnail may correspond to a representative frame of video segment a in "Video A". The representative frame may be the starting frame, ending frame, intermediate frame, optimal frame, or any other frame of video segment a. The starting frame of video segment a may also be used as the thumbnail, which is not limited in this application.

[0309] For example, the first time point can be the same as the time point of the video frame corresponding to the thumbnail. For example, if the thumbnail corresponds to the representative frame of video segment a, the first time point can be the time point corresponding to the representative frame; the first time point can also be different from the time point corresponding to the thumbnail. For example, the thumbnail is the representative frame of video segment a, the representative frame is the middle frame of video segment a, and the first time point can be the time point corresponding to the starting frame of the video segment a. This application does not limit this.

[0310] It should be noted that the above-mentioned method of displaying the search box 321 through the album display interface 320 for the user to input search text is only an example, and the search box can also be displayed through other interfaces of the gallery.

[0311] In some embodiments, as shown in FIG3a , the album display interface 320 has a “Photo” control, a “Time Point” control, and a “Create” control at the bottom, which are used to enter the photo display interface, the time point display interface, and the creation display interface of the gallery, respectively. The photo display interface, the time point display interface, and the creation display interface all include a search box. The search boxes of these three interfaces support the same search function as the search box 321 of the album display interface 320 , and all support searching for pictures and videos included in the gallery based on the search text entered by the user. After the search is completed, the first search result display interface 330 is also displayed, and the first search result included is also the same.

[0312] It should be understood that if the vector similarity between the visual semantic vector of the representative frame of the video segment and the text semantic vector of a certain text exceeds a certain threshold, the video segment can be searched when the user inputs the text.

[0313] In some embodiments, a video includes multiple video segments, and the search text entered by the user may match an index of one or more video segments of a video.

[0314] For example, the search text "boy standing outdoors" entered by the user can match video segments a, b, and c in "Video A". Therefore, video segments a, b, and c of "Video A" can be displayed simultaneously. The thumbnail displayed for video segment b of "Video A" can be the representative frame of video segment b, and the thumbnail displayed for video segment c of "Video A" can be the starting frame of video segment c. As shown in the first display area 331 of the first search result display interface 330 of Figure 3a, the search text "boy standing outdoors" entered by the user can search for video segments a, b, and c of Video A. Video segments a, b, and c are sorted in order of their matching degree with the search text. This indicates that the vector similarity between the visual semantic vectors of the representative frames of these three video segments and the text semantic vector of "boy standing outdoors" exceeds the vector similarity threshold.

[0315] In other embodiments, different search texts input by the user may all search for the same video segment.

[0316] As shown in (1) of FIG3b , the first search result display interface 350 may include a first display area 351 and a second display area 352. The first display area 351 displays all video search results. The second display area 352 displays all image search results, and the right side of each image search result may display the shooting time and shooting location of the image.

[0317] Assume that the user enters "boy standing by the river" in the input box, and the first display area 351 can display video search results that match it, including video segment b of "Video B". The first display area 351 includes a thumbnail of the representative frame of video segment b and two time points. The first time point is the time "05:48" corresponding to the starting frame of video segment b, and the second time point represents the total duration of "Video B" "14:18". If the user triggers the thumbnail of the representative frame of video segment b, the mobile phone 300 can display the video playback interface 360, in which the mobile phone 300 adjusts the time progress of "Video B" to "05:48" to start playing, that is, the mobile phone 300 starts displaying from the starting frame of video segment b in "Video B".

[0318] Optionally, the first time point may also be the time point of a representative frame of video segment b of video B, which is not limited in this solution.

[0319] It should be noted that (1) in FIG3b is only an example. The first display area 351 can also display other information of "Video B", such as the shooting location, shooting time, names and relationships of people corresponding to "Video B". The second display area 352 can also display other information of the image search results, such as the names and relationships of people corresponding to the image.

[0320] It should be noted that after the user triggers the thumbnail of the representative frame of video segment b, mobile phone 300 adjusts the time progress of "Video B" to the first time point displayed for playback. It should be understood that when the first time point is the time point of the representative frame of video segment b of video B, if the middle frame of video segment b is the representative frame, the time progress can also be adjusted to the time point corresponding to the middle frame to start playback. The time point corresponding to the middle frame is "07:09", which is later than the time "05:48" corresponding to the starting frame.

[0321] As shown in (2) of FIG3b , the first search result display interface 370 may include a first display area 371 and a second display area 372. The first display area 371 displays all image search results, and the second display area 372 displays all video search results.

[0322] Assuming that the user enters "boy standing near the bridge" in the input box, the second display area 372 can display the video search results that match it, including the video segment c of "video B" and the video segment b of video "B". Combined with (1) in Figure 3b, the two different search texts "boy standing by the river" and "boy standing near the bridge" entered by the user both searched for the video segment b of video "B". This shows that the visual semantic vector of the representative frame of the video segment b of "video B" has a vector similarity exceeding the vector similarity threshold with the text semantic vector of "boy standing by the river", and also exceeds the vector similarity threshold with the text semantic vector of "boy standing near the bridge". The display form of the video segment b of "video B" can be found in the introduction of (1) in Figure 3b, which will not be repeated here.

[0323] The first display area 372 includes a thumbnail of the representative frame of the video segment c and two time points. The first time point is the time "08:32" corresponding to the representative frame of the video segment c, and the second time point represents the total duration of "Video B" "14:18". If the user triggers the thumbnail of the representative frame of the video segment c, the mobile phone 300 can display the video playback interface 380. In this interface, the mobile phone 300 adjusts the time progress of "Video B" to "07:26" to start playing, that is, the mobile phone 300 starts playing from the starting frame of the video segment b in "Video B". The display format of the video search results (2) and the image search results in Figure 3b can be found in the introduction of Figure 3a and Figure 3b (1), which will not be repeated here.

[0324] It should be understood that recommended information can be displayed in the search box of the gallery application, which is helpful for users to obtain relevant information about the pictures / videos stored in the gallery.

[0325] In some embodiments, the recommended information displayed in the search box of the gallery is a sentence with natural semantics. The user can enter text in the search box based on the content and format of the recommended information, using the sentence with natural semantics as the search text entered into the search box. The electronic device searches for images / videos in the gallery based on the search text with natural semantics, which can more accurately search for the images / videos the user needs.

[0326] In one implementation, the electronic device may obtain attribute tags of the picture / video. In one example, the attribute tags include at least one of time, location, category tag, and event.

[0327] In some embodiments, the “time” can be obtained based on the shooting time of the picture / video. For example, the “time” may include “Thursday, November 9, 2023”, “Friday, July 2, 2021”, etc.

[0328] The "location" can be obtained based on the location where the picture / video was taken. For example, it can be obtained based on the GPS positioning information when the video was taken. For example, the "location" can include a city, a scenic spot, etc.

[0329] "Classification tags" and "events" can be obtained through semantic analysis of images / videos. In one example, the electronic device can use computer vision services to perform semantic analysis on the video frames in the image or video, and generate "classification tags" and "event" content based on the semantics.

[0330] For example, "category tags" may include: people, plants or animals, etc., or may include: architecture, natural scenery, etc., or may include: street art, musical instruments, art exhibitions, competitions, birthdays, etc. For example, a person may include the name of the portrait, the name of the occupation, the name of the person's age group, etc.

[0331] For example, “events” may include: games, sports, travel, etc.

[0332] It should be understood that for a picture / video, its corresponding attribute tag may include one or more of time, location, classification tag, and event. For example, for a picture downloaded from the Internet, if the electronic device does not obtain its shooting time, its attribute tag does not include "time" or the content of "time" is empty.

[0333] In one implementation, one or more attribute tags of an image / video are combined according to preset rules to generate at least one piece of combined information. For example, the electronic device combines the four attributes of a video: "time," "location," "category tag," and "event" to generate a piece of combined information. In another example, the electronic device can also generate a piece of combined information based on a single attribute tag of the image / video, such as "time."

[0334] In one implementation, fixed splicing words are spliced ​​with a combination of information to generate recommended content corresponding to the picture / video. For example, the fixed splicing words may include "try to search", "at", "shoot", "video of", etc.

[0335] Table 1 shows some examples of recommended content generated based on different numbers of attribute tags. It should be noted that when generating recommended content, the attribute tags can be mapped accordingly. For example, the specific time "Monday, October 1, 2023" can be mapped to "today," "the day before yesterday," "August," "last month," "last year," "this year," or "National Day."

[0336] Table 1

[0337] As shown in Table 1, a video's attribute tags can be combined to generate multiple spliced ​​contents according to the splicing combination method shown in Table 1. For example, if a video's attribute tags include "category tag," "time," and "location," five different spliced ​​contents can be generated according to the splicing combination methods numbered 1-5 in Table 1.

[0338] In one implementation, the electronic device may determine any one of the multiple spliced ​​contents as the recommended content corresponding to the picture / video.

[0339] In another implementation, each of the above-mentioned splicing combination methods corresponds to a priority level, and the priorities of the splicing combination methods decrease in order from front to back in Table 1. The electronic device generates the recommended content corresponding to the picture / video according to the splicing combination method with the highest priority among the splicing combination methods supported by the picture / video. As shown in Table 1, if the attribute tags of the video include "classification tag", "time", and "location", the fixed splicing words are spliced ​​with the "classification tag", "time", and "location" according to the splicing combination method of sequence number 1 in Table 1 to generate the recommended content corresponding to the picture / video.

[0340] In the visual media search method provided in the embodiment of the present application, recommendation information in the search box can be generated based on the recommended content corresponding to a picture / video in the gallery.

[0341] In one implementation, the electronic device may select a recommended content corresponding to a first image / video in the gallery to generate recommendation information in the search box. In one example, the recommended content may be a first image / video within a preset time period; for example, the preset time period is more than one month from today. In another example, the electronic device updates the first image / video once a day, and the first image / video selected within a preset time period (e.g., within a week) is not repeated.

[0342] For example, on the first day, the recommendation displayed in the search box in the electronic device's gallery is "Try searching for videos of Zhang San walking by the river the day before yesterday." On the second day, the recommendation displayed in the search box in the electronic device's gallery is "Try searching for pictures of the beach last August." ... Within a preset period (e.g., a week), the first picture / video selected by the electronic device each day is unique, and accordingly, the recommendation displayed in the search box in the electronic device's gallery is unique each day.

[0343] Based on the scenario example shown in FIG3a , still taking the electronic device as a mobile phone 300, assume that a user takes multiple videos on a weekend trip using the mobile phone 300 and stores them in the photo gallery of the mobile phone 300. The mobile phone can then generate recommendation information in the search box based on the recommended content corresponding to the videos.

[0344] As shown in Figure 3c, mobile phone 300 displays an album display interface 390, which includes a search box 321. In one example, search box 321 displays a recommendation message, "Try searching for videos of Zhang San by the river last year." This recommendation message serves as an example of the search text entered by the user. The user can enter the search text in the search box based on the content and format of the recommendation message to conduct a search.

[0345] As shown in FIG3c , after a user clicks to trigger search box 321, mobile phone 300 displays search interface 301, which includes search box 321, into which the user can enter text. Upon detecting the user entering text in search box 321, mobile phone 300 can search the gallery based on the search text in search box 321.

[0346] For example, the search box 321 displays the recommended information "Try searching for the video of Zhang San by the river last year." "Zhang San by the river last year" is a sentence with natural semantics. Compared with single tags such as "last year," "by the river," and "Zhang San," it is easier to search for the video that the user really needs to find. For example, in the gallery, there are 7,230 pictures / videos corresponding to the tag "last year," 1,985 pictures / videos corresponding to the tag "by the river," and 1,985 pictures / videos corresponding to the tag "Zhang San." The number of pictures / videos searched using a single tag is relatively large. However, when searching in the gallery based on "Zhang San by the river last year," the number of pictures / videos searched is significantly reduced. In this way, displaying sentences with natural semantics as recommended information is conducive to providing users with a convenient and fast search experience.

[0347] As another example, the visual media search method provided in the embodiments of the present application can be implemented in a global search scenario of an electronic device.

[0348] In conjunction with FIG4a , the following still uses a mobile phone 300 as an example electronic device to exemplify the visual media search method provided by an embodiment of the present application. In this scenario, the image library of mobile phone 300 may include multiple images and multiple videos. The sources of the images / videos can be found in the above examples and will not be further described here.

[0349] In conjunction with Figure 4a , based on the example scenario shown in Figure 3a above, let's assume that a user takes multiple videos on their phone 300 during a weekend trip and stores them in the photo gallery of their phone 300. As shown in Figure 4a , the phone 300 displays a main interface 410. The user can swipe right on the main interface 410 to display a negative one-screen interface 420. This negative one-screen interface 420 provides a global search function, offering users a wide range of online search services, as well as search services for the phone's local resources (i.e., local files stored on the phone 300).

[0350] Negative-one screen interface 420 includes a search box 421. A user can enter the search text "boys standing outdoors" in search box 421. Mobile phone 300 can then perform an online search based on the user-entered search text, as well as a local resource search within local files stored on mobile phone 300. After the search is complete, mobile phone 300 displays a second search result display interface 430, which can display a second search result corresponding to the input "boys standing outdoors."

[0351] As shown in Figure 4a, the second search result display interface 430 may include a first display area 431, a second display area 432, and a third display area 433. In some embodiments, the first display area 431 includes the second search result of the online search, the second display area 432 includes the second search result of the application search, and the third display area 433 includes the second search result of the local file search.

[0352] In some embodiments, the second search results for the application include the second search results for the gallery application. For example, the content displayed in second display area 432 is the content displayed in the portion of first search results display interface 330 in Figure 3a. The second search results included in second display area 432 are pictures and videos stored in the gallery. The thumbnails and time points displayed for the videos can be found in the above examples and are not further described here.

[0353] It should be noted that the content displayed in the second display area 432 is only an example, and the content displayed in other areas of the first search result display interface 330 in FIG. 3 a may also be displayed.

[0354] It should be noted that the above-mentioned method of displaying the search box 421 for the user to enter search text to perform a global search through the negative first screen interface 420 is merely an example. In some embodiments, the user can also pull down the main interface of the mobile phone 300 to display the main menu interface of the mobile phone 300, which includes a search box and also provides a global search function. The search results displayed are also the same as the second search results.

[0355] Also illustratively, the visual media search method provided in the embodiments of the present application can be implemented in a search scenario of a video playback application of an electronic device.

[0356] In this scenario, the visual media search method is implemented through the interaction between the electronic device and the cloud server.

[0357] Still taking the electronic device as a mobile phone 300 as an example, in this scenario, the mobile phone 300 includes a video playback application.

[0358] As shown in Figure 4b, mobile phone 300 displays a main interface 440, which includes application icons for multiple applications, including an application icon 441 for a video playback application. A user triggers application icon 441 for the video playback application, and mobile phone 300 displays a first interface 450 for the video playback application. This first interface 450 includes a search box 451 and a search control 452. The user can enter the search text "documentary about city B" in search box 451 and trigger search control 452. In response to the user's triggering operation on search control 452, mobile phone 300 searches multiple videos on the cloud server based on the search text entered by the user. After the search is complete, mobile phone 300 displays a third search result display interface 460 for the video playback application. This search result display interface includes multiple third search results, which are videos. The thumbnails and time points displayed can be found in the above example and will not be repeated here.

[0359] In some embodiments, the third search result display interface 460 may also display the publication time of the third search result, i.e., the time when the third search result was stored on the cloud server. As shown in FIG4b , taking "Video C" included in the third search result display interface 460 as an example, the third search result display interface 460 also displays the publication time of "2018-05-03," indicating that "Video C" was stored on the cloud server on May 3, 2018. For example, the user may have uploaded "Video C" to the video playback application on May 3, 2018, causing the cloud server to store "Video C."

[0360] In some embodiments, as shown in FIG4b , the third search result display interface 460 includes "Video C," where the first time point of "Video C" is "05:48." If "Video C" is the video the user desires, the user can trigger "Video C," and the mobile phone 300 can interact with the cloud server to start playing "Video B" from "05:48."

[0361] Next, let’s introduce the second application scenario based on person relationship search.

[0362] In addition, if multiple video segments are hit in the video, one of the multiple hit video segments can be selected and the information of the video segment can be displayed on the video cover in the above manner. In one embodiment, the multiple hit video segments can be sorted according to the degree of match with the search text, and the information of the hit video segment with the highest degree of match with the search text can be displayed on the video cover in the above manner.

[0363] The method provided in the embodiment can display search results in multiple ways, and the embodiments of the present application do not impose any limitations on this.

[0364] Next, the effect interface of the method provided in the embodiment of the present application is described by taking the search text including character relationships as an example.

[0365] Exemplarily, FIG4c is a schematic diagram of the interface changes of another example of a visual media search method provided in an embodiment of the present application. As shown in FIG4c (a), the user enters the search text "traveling with family last year" in the search box 103 in the album interface 102 of the mobile phone. Among them, "family" belongs to character relationship information. In response to the user's operation, the mobile phone searches the gallery database for images and videos that match "traveling with family last year", and displays the search results in the search result interface 601, as shown in FIG4c (b). It can be seen that in the embodiment of the present application, when the search text contains character relationships, images and / or videos that meet the character relationships can also be searched, which meets the user's expectations and improves the user experience.

[0366] In this embodiment, the display method of the images and videos in the search results may be the same as that in the embodiment shown in FIG. 3 a above, and will not be described in detail.

[0367] As shown in FIG4c(b), for example, if the search results include video 602, the user can click on video 602. In response to the user's operation, the mobile phone jumps to video playback interface 603, as shown in FIG4c(c). Video playback interface 603 includes a playback control 604. The user clicks on play control 604, and the mobile phone begins playing video 602 in response to the user's operation, as shown in FIG4c(d).

[0368] In one embodiment, the mobile phone may start playing the video from the start time of the hit video segment, as shown in FIG4c (d), where the video 602 starts playing from 2 minutes and 26 seconds (02:26).

[0369] In another embodiment, the video may also be played starting from the play time corresponding to the representative frame of the hit video segment.

[0370] In yet another embodiment, the video may also be played starting from the start time (00:00) of the video 602 .

[0371] In another embodiment, when video 602 includes multiple matching videos, the matching videos may be played sequentially according to their chronological order in video 602, or may be played sequentially according to their degree of matching with the search text. This embodiment of the present application does not impose any limitation on this.

[0372] As will be appreciated, FIG4c illustrates the effectiveness of the visual media search method using the example of a search text containing the character relationship information of "family." It will be appreciated that the search text may also include other character relationships, such as "colleague," "friend," "classmate," "son," "daughter," "best friend," and so on.

[0373] Then we will introduce the third application scenario of search result sorting.

[0374] For example, Figure 4d is a schematic diagram of the interface changes of an example visual media search method provided in an embodiment of the present application. As shown in Figure 4d (a), the mobile phone screen displays the album interface 102 of the Gallery app. Album interface 102 includes a search box 103. The user enters the search text "boy standing outdoors" in search box 103. In response to the user's operation, the mobile phone searches the gallery database for images and videos matching "boy standing outdoors" and displays the search results, as shown in the search results interface 501 of Figure 4d (b).

[0375] In one embodiment, when displaying search results, the videos and images in the search results can also be sorted according to the degree of matching with the search text. Optionally, the sorting principle can be: the higher the degree of matching with the search text, the higher the ranking when displayed. Among them, the degree of matching between the image or video and the search text can be represented by a matching score. The higher the matching score, the higher the degree of matching. For example, in Figure (b) of Figure 4d, the search results include video 5051, image 5052, video 5053, image 5054, image 5055, image 5056, image 5057 and image 5058 from left to right and from top to bottom, and the matching scores of these videos or images decrease in sequence.

[0376] The degree of match between a video and the search text can be represented by the highest matching score for the video segment. For example, in Figure 4d (b), video 5051 includes three matching segments, of which the first matching segment has the highest matching score. This matching score is then used as the score for video 5051 and compared with the scores of other videos or images for ranking.

[0377] In addition, it can be understood that when the user clicks the "more" control 506, the electronic device displays all search results in response to the user's operation. All search results can also be sorted and displayed according to the degree of matching with the search text, which will not be repeated here.

[0378] The following describes how videos are displayed in search results.

[0379] In one embodiment, the video cover may display relevant information about the video. For example, as shown in FIG4d (b), the cover of video 5051 may display the total duration of the video at 8 minutes and 32 seconds (08:32). The cover of video 5053 may display the total duration of the video at 14 minutes and 18 seconds (14:18).

[0380] In one embodiment, the cover of the video in the search results may further display the start time (referred to as the start time), end time (referred to as the end time), and middle time (referred to as the middle time) of a particular hit video segment. For example, if a video includes three hit video segments, the video cover may display the start time of the first hit video segment.

[0381] In another embodiment, the start time, end time, or middle time of the highest matching video segment can also be displayed on the video cover. As shown in FIG4d (b), taking the start time of the highest matching video segment in video 5051 as 2 minutes and 18 seconds (02:18) and the end time as 3 minutes and 18 seconds (03:18) as an example, the cover of video 5051 can display the start time as 2 minutes and 18 seconds (02:18), the end time as 3 minutes and 18 seconds (03:18), or the middle time as 2 minutes and 48 seconds (02:48).

[0382] In another embodiment, after a video is divided into video segments, a frame of image can be selected from each video segment as a representative frame, and the video cover can also display the corresponding playback time of the representative frame of the matching video segment. For example, if the corresponding playback time of the representative frame of the matching video segment is 2 minutes and 0 seconds (02:00), the video cover can display this time.

[0383] In one embodiment, the cover of a video can be a thumbnail of any frame image in the video. For example, it can be the first frame image of the video, or the first frame image of the hit video segment with the highest matching score, or the middle frame of the hit video segment with the highest matching score, or the representative frame of the hit video segment with the highest matching score, etc. For example, as shown in (b) of FIG4d , the cover image of video 5051 is the first frame image of the hit video segment with the highest matching score.

[0384] The above embodiments are all described by taking the video cover in the search results showing the information of one hit video segment as an example. In other embodiments, the video cover in the search results can also show the information of multiple hit video segments.

[0385] For example, FIG4e is a schematic diagram of a search result interface provided by an embodiment of the present application. As shown in FIG4e, the video cover can display the total number of hit video segments and the total number of video segments contained in the video, so that the user can clearly know the number of hit video segments in the video. For example, video 5051 is divided into 5 video segments, 3 of which match the search phrase "boy standing outdoors", that is, it includes 3 hit video segments, then the cover of video 5051 can display information such as "3 hit segments / 5 segments in total". The same is true for video 5053, which will not be repeated here.

[0386] In another embodiment, the video cover may further display one or more of the following information, such as the start time, end time, intermediate time, or playback time corresponding to a representative frame of the multiple hit video segments, so that the user can clearly know the location of the hit video segments in the video. For example, the start time of the multiple hit video segments may be displayed, or the playback time corresponding to the representative frames of the multiple hit video segments may be displayed.

[0387] Assume that the three hit video segments of video 5051 are the first, third, and fifth video segments. The first video segment starts at 2 minutes and 0 seconds (02:18) and ends at 3 minutes and 18 seconds (03:18), and the corresponding playback time of the representative frame is 2 minutes and 10 seconds (02:10); the second video segment starts at 5 minutes and 22 seconds (05:22) and ends at 6 minutes and 40 seconds (06:40), and the corresponding playback time of the representative frame is 5 minutes and 45 seconds (05:40); the third video segment starts at 8 minutes and 0 seconds (08:00) and ends at 8 minutes and 32 seconds (08:32), and the corresponding playback time of the representative frame is 8 minutes and 12 seconds (08:12).

[0388] For example, Figure 4f is a schematic diagram of another search result interface provided by an embodiment of the present application. Optionally, referring to Figure 4f , the cover of video 5051 can display information such as "Hit 02:18-03:18," "Hit 05:22-06:40," and "Hit 08:00-08:32." The same applies to video 5053, which will not be further described.

[0389] For example, Figure 4g is another example of a search result interface diagram provided in an embodiment of the present application. Optionally, referring to Figure 4g , the cover of video 5051 can display information such as "Hit 02:10," "Hit 05:40," and "Hit 08:12." The same is true for video 5053, which will not be further described.

[0390] It is understandable that when there are a large number of matching video segments, the video cover can display information of N video segments with higher matching scores, where N can be 3, 2, etc. In this way, it can prevent the information from blocking too much of the video cover and affecting the user experience.

[0391] In one embodiment, the cover of the video may also display frame images from multiple hit video segments. For example, the representative frames of N video segments with higher matching scores may be spliced ​​into one image as the cover of the video, so that the user can intuitively see the content in the hit video segments and improve the user experience. For example, FIG4h is another schematic diagram of a search result interface provided in an embodiment of the present application. Referring to FIG4h, the representative frames of the first hit video segment, the second hit video segment, and the third hit video segment of video 5051 may be spliced ​​together to serve as the cover of video 5051. The same is true for video 5053 and will not be described in detail.

[0392] In summary, the method provided in the embodiments of the present application can have multiple ways of displaying search results, and the embodiments of the present application do not impose any restrictions on this.

[0393] The following describes the process of playing videos in the search results.

[0394] For example, Figure 4i is a schematic diagram of an interface illustrating a video playback process provided in an embodiment of the present application. As shown in Figure 4i (a), a user clicks on video 5051 in the search results. In response to the user's operation, video playback interface 1001 is displayed. Video 5051 and playback controls 1002 are displayed in video playback interface 1001. The cover of video 5051 can be consistent with the cover display effect described in the above embodiments.

[0395] The user clicks the play control 1002, and in response to the user's operation, the mobile phone begins playing video 5051. Optionally, the mobile phone can play each hit video segment one at a time in the order of the playback time period, that is, first play the first hit video segment (02:18-03:18), then play the second hit video segment (05:22-06:40), and then play the third hit video segment (08:00-08:32). After the three hit video segments are played, playback stops. Taking the image displayed on the video cover as the first frame image of the first hit video segment as an example, after the user clicks the play control 1002, the playback interface of video 5051 can be shown as shown in Figure 4i (b). It can be seen that after the video starts playing, a playback progress bar 1003 can be displayed in the interface. Currently, the video starts playing from the first hit video segment, so the playback time displayed in the playback progress bar 1003 is 02:18.

[0396] Optionally, the mobile phone can also play the matching video segments in descending order of their matching scores. For example, if the matching score of the second matching video segment is greater than that of the first matching video segment, and greater than that of the third matching video segment, the matching video segments will be played in the order of the second matching video segment, the first matching video segment, and the third matching video segment. Playback stops after all three matching video segments have finished playing.

[0397] Furthermore, it is understood that when playing each of the hit video segments, playback can start from the first frame of the hit video segment, from an intermediate frame of the hit video segment, or from a representative frame of the hit video segment. The representative frame best represents the content of the video segment. Therefore, starting playback of the hit video segment from the representative frame allows users to quickly and directly view the scenes in the video segment related to the search text, thereby improving the user experience.

[0398] Exemplarily, FIG4j is an interface diagram of another example of a video playback process provided by an embodiment of the present application. In some embodiments, when a video is played, information about each hit video segment may be displayed in the playback progress bar 1003. For example, the position of each hit video segment in the video is displayed in the playback progress bar 1003, as shown in FIG4j (a), where the hit video segment is shown as a black long bar. Optionally, corresponding text may also be displayed above the playback progress bar 1003 to indicate the hit video segment. For example, the numbers "1", "2", and "3" are displayed to indicate the three hit video segments; or according to the matching score, "match ranking 1", "match ranking 2", "match ranking 3", etc. are displayed to indicate the matching ranking of the video segment and the search text.

[0399] In another embodiment, the matching scores of the various hit video segments can be distinguished by different colors. For example, the hit video segment with the highest matching score is marked in red in the progress bar, the hit video segment with the second highest matching score is marked in pink in the progress bar, and the hit video segment with the lowest matching score is marked in green in the progress bar, as shown in Figure 4j (b). It should be noted that in Figure 4j (b), different colors are represented by different fill patterns, which is only an example and does not represent the actual display effect.

[0400] Exemplarily, Figure 4k is an interface diagram of another example of a video playback process provided by an embodiment of the present application. Optionally, marking information, such as a video tag, may also be provided at the hit video segment. As shown in Figure (a) in Figure 4k, the video tag is shown as a black inverted triangle. The user clicks on the tag, and the video jumps directly to the corresponding hit video segment to start playing. As shown in Figure (a) in Figure 4k, when the video plays to 02:25, the user clicks on the video tag of the second hit video segment (05:22-06:40) with the highest matching score, and the hit video segment starts playing, as shown in Figure (b) in Figure 4k.

[0401] In this implementation, the hit video segment and / or its matching ranking are displayed in the video's playback progress bar, allowing users to intuitively identify the location of content related to the search text in the video, improving the user experience. Furthermore, video tags are inserted at the hit video segment in the video. Users can click on the video tag to jump directly to the corresponding hit video segment, eliminating the need to drag the playback progress bar to find relevant content. This facilitates user operation and further improves the user experience.

[0402] The above embodiment is explained by taking the search in the album interface 102 in the gallery as an example. In actual applications, the search can also be implemented in other interfaces of the gallery APP, or in other application or service interfaces. For example, referring to Figure (a) in Figure 41, you can enter content in the search box 1302 in the photo interface 1301 of the gallery to implement the search. For another example, referring to Figure (b) in Figure 41, you can enter content in the search box 1304 in the moment interface 1303 of the gallery to implement the search. For another example, referring to Figure (c) in Figure 41, you can also enter content in the global search box 1306 in the negative one screen interface 1305 of the mobile phone to implement the search. The embodiments of the present application do not impose any restrictions on the search interface and search entry.

[0403] Then the fourth application scenario of automatic search is introduced.

[0404] The present invention provides a visual media search method that supports a user inputting a natural-sense sentence and searches for image resources in a gallery based on the natural-sense sentence. The electronic device displays recommended information to the user in a gallery search box. The recommended information serves as an example of the user's input content, prompting the user to enter text information in the search box in terms of content and format.

[0405] Still taking mobile phone 100 as an example, as shown in FIG4A , mobile phone 100 displays a photo display interface 401. Photo display interface 401 includes a search box 402 and a photo display page 403. Search box 402 is used to trigger the display of an image resource search interface, and photo display page 403 is used to display thumbnails of image resources (pictures or videos) in the gallery. In one implementation, as shown in FIG4A , the image resource thumbnails in photo display page 403 are arranged from recent to remote according to the time of the image resources.

[0406] In one example, the search box 402 displays a recommendation message, "Try searching for photos of a beach walk last August." This recommendation message serves as an example of user input content. The user can enter text information in the search box to search based on the content and format of the recommendation message.

[0407] For example, as shown in FIG4A , in response to a user clicking on search box 402, mobile phone 100 displays search interface 201, which includes search box 202, into which the user can enter text. Upon detecting the user entering text in search box 202, mobile phone 100 can search the gallery based on the text in search box 202. For example, search box 202 displays a recommendation message: "Try searching for photos of a beach walk last August."

[0408] Among them, "Walking on the beach last August" is a sentence with natural semantics. Compared with single tags such as "beach" or "dinner party", it is more likely to hit the image resources that users are really looking for. For example, in the gallery, there are 7,230 image resources corresponding to the tag "last year", 2,651 image resources corresponding to the tag "August", and 1,985 image resources corresponding to the tag "beach". The number of image resources searched using a single tag is relatively large. However, searching the gallery based on "Walking on the beach last August" will significantly reduce the number of image resources found.

[0409] For example, as shown in FIG4B , a user enters the text "last August beach walk" in search box 202 and then clicks search box 202. Mobile phone 100 receives the user's input of "last August beach walk" in search box 202 and searches the image library for image resources that are semantically consistent with "last August beach walk." As shown in FIG4B , mobile phone 100 displays search results interface 501, which displays image resources found in the image library based on the text in search box 202. For example, 54 image resources in the image library that are semantically consistent with "last August beach walk" are found.

[0410] In the visual media search method provided by the embodiments of this application, the recommended information displayed in the search box of the image gallery is a sentence with natural semantics. The user can enter text information in the search box based on the content and format of the recommended information, using the sentence with natural semantics as the input content of the search box. The mobile phone searches for image resources in the image gallery based on the text information with natural semantics, and can more accurately hit the pictures and videos that the user is searching for.

[0411] In one implementation, the mobile phone may obtain constituent elements of the image resource. In one example, the constituent elements include at least one of time, location, person, subject type, and event.

[0412] The constituent element "time" can be obtained according to the shooting time of the image resource; for example, the constituent element "time" includes "Thursday, November 9, 2023", "Friday, July 2, 2021", etc.

[0413] The constituent element "location" can be obtained based on the shooting location of the image resource, such as based on the GPS positioning information when the image resource is shot; for example, the constituent element "location" can include a city (such as Beijing) and / or a scenic spot name (such as the Great Wall).

[0414] The elements "person," "subject type," and "event" can be obtained through semantic analysis of image resources. For example, a mobile phone can use a semantic analysis algorithm to perform semantic analysis on images in pictures or videos and generate content for "person," "subject type," and "event" based on the semantics of the image resources.

[0415] For example, the element “character” may include:

[0416] Names of people (such as Xiaohua, Xiaoming, Zhang San, Li Si), acrobats, orchestras, young people, children, babies, etc.

[0417] For example, the element "subject type" may include:

[0418] Musical instruments, calligraphy, paintings, animals, plants, people, scenery, beaches, outdoors, architecture, birthdays, etc.

[0419] For example, the component "event" may include:

[0420] Games, singing, dancing, sports, barbecues, walks, travel, etc.

[0421] It's understandable that for an image resource, its corresponding constituent elements may include one or more of time, location, person, subject type, and event. For example, for an image downloaded from the internet, if the phone doesn't know when it was taken, its constituent elements won't include "time" or will be empty. For example, for a landscape image, if semantic analysis doesn't find any "person" content, its constituent elements won't include "person" or will be empty.

[0422] In one implementation, at least one piece of combination information may be generated by splicing constituent elements of an image resource according to a preset rule.

[0423] In one example, a mobile phone combines the four elements of an image resource, namely, "time," "place," "person," and "event," to generate a combined message.

[0424] In one example, the mobile phone combines the three elements of an image resource to generate a combined piece of information. For example, "place," "person," and "event" are combined; "time," "place," and "person" are combined; "time," "place," and "event" are combined; "time," "person," and "event" are combined; "time," "place," and "subject type" are combined; and so on.

[0425] In one example, the mobile phone combines two elements of an image resource to generate a combined piece of information. For example, "person" and "event" are combined, "location" and "event" are combined, "time" and "event" are combined, "location" and "person", "time" and "person", "location" and "subject type", "time" and "location", etc.

[0426] In one example, the mobile phone may also generate a piece of combined information based on a constituent element of the image resource, for example, the constituent element is "time".

[0427] In one implementation, fixed concatenated words are concatenated with a combination of information to generate recommended content corresponding to the image resource. The fixed concatenated words include "try to search", "at", "shoot", "photo of", etc.

[0428] Table 2 shows some examples of recommended content generated based on different numbers of constituent elements. It should be noted that when generating recommended content, the constituent elements can be mapped accordingly. For example, the specific time "Monday, August 7, 2023" can be mapped to "today," "the day before yesterday," "August," "last month," "last year," or "this year."

[0429] Table 2

[0430] The constituent elements of an image resource can be combined and combined according to the splicing methods shown in Table 2 to generate multiple splicing contents. For example, if the constituent elements of an image resource include "person," "time," "place," and "event," 14 different splicing contents can be generated according to the splicing methods numbered 1-14 in Table 2. If the constituent elements of an image resource include "place," "person," and "event," 4 different splicing contents can be generated according to the splicing methods numbered 2, 7, 8, and 10 in Table 2.

[0431] In one implementation, the mobile phone determines any one of the multiple spliced ​​contents as the recommended content corresponding to the image resource.

[0432] In another implementation, each of the above-mentioned splicing combination methods corresponds to a priority level, and the priorities of the splicing combination methods decrease in order from front to back in Table 2. The mobile phone generates the recommended content corresponding to the image resource according to the splicing combination method with the highest priority among the splicing combination methods supported by the image resource. For example, if the constituent elements of an image resource include "place", "person" and "event" but not "time", the fixed splicing words are spliced ​​with "place", "person" and "event" according to the splicing combination method of sequence number 2 in Table 2 to generate the recommended content corresponding to the image resource.

[0433] In the visual media search method provided in the embodiments of the present application, recommendation information within the search box is generated based on the recommended content corresponding to an image resource in the image library. For example, if the recommended content corresponding to an image resource is "Try searching for photos of Xiaohua" + "the day before yesterday" + "at" + "Big Wild Goose Pagoda" + "travel" + photos, the recommendation information generated based on this recommended content is "Try searching for photos of Xiaohua traveling at the Big Wild Goose Pagoda the day before yesterday."

[0434] In one implementation, the mobile phone selects recommended content corresponding to a first image resource in the gallery to generate recommendation information within the search box. In one example, the first image resource is an image resource within a preset time period; for example, the preset time period is more than one month from today. In another example, the mobile phone updates the first image resource once a day, and the first image resource selected within the preset time period (e.g., within a week) is not repeated.

[0435] For example, on the first day, the search box in the phone's gallery displays the recommendation "Try searching for photos of Xiaohua traveling to the Big Wild Goose Pagoda the day before yesterday." On the second day, the search box displays the recommendation "Try searching for photos of a beach walk last August." On the third day, the search box displays the recommendation "Try searching for photos of animals taken at the zoo." Within a preset time period (e.g., a week), the phone's selected first image resource each day remains unique, and accordingly, the recommended information displayed in the search box in the phone's gallery remains unique each day.

[0436] It should be noted that the visual media search method provided in the embodiments of the present application can be applied to the search box in any user interface of a gallery. Figure 4A illustrates the example of a photo display interface 401 including a search box 402, and does not limit the application scenarios of the embodiments of the present application. For example, as shown in Figure 4C, mobile phone 100 displays a desktop interface including a gallery application icon 101. In response to a user clicking on application icon 101, mobile phone 100 displays a gallery album display interface 102, which includes a search box 103. Search box 103 displays a recommendation message, "Try searching for photos of a beach walk last August." This recommendation message serves as an example of text information entered by the user. The user can enter text information in the search box to search based on the content and format of the recommendation message. For example, in response to a user clicking on search box 103, mobile phone 100 displays a search interface 201, which includes a search box 202, which displays a recommendation message, "Try searching for photos of a beach walk last August."

[0437] Referring to Figure 4B , a user enters the text "Walking on the beach last August" in search box 202 and then clicks search box 202. Mobile phone 100 receives the user's input of "Walking on the beach last August" in search box 202 and searches the gallery for image resources that are semantically consistent with "Walking on the beach last August." As shown in Figure 4B , mobile phone 100 displays search results interface 501, which displays image resources found in the gallery based on the text in search box 202. For example, 54 image resources in the gallery that are semantically consistent with "Walking on the beach last August" are found.

[0438] In some embodiments, the mobile phone pre-analyzes the image resources stored in the gallery, obtains recommended content corresponding to each image resource in the gallery, and stores the recommended content corresponding to each image resource in the gallery. For example, when the processor of the mobile phone is idle, the mobile phone performs data analysis on the image resources stored in the gallery to generate recommended content corresponding to each image resource in the gallery. For example, each time an image resource in the gallery changes (e.g., is added, deleted, or edited), the mobile phone performs data analysis on the image resources stored in the gallery to generate recommended content corresponding to each image resource in the gallery. Before the search box is displayed (e.g., in the scenario shown in FIG4C , after the mobile phone 100 receives a user click on the application icon 101), the mobile phone selects a first image resource in the gallery based on preset conditions (e.g., the image resource must have been taken at least one month ago; for example, the recommended content must not be repeated within a week), and then obtains recommended content corresponding to the first image resource from the recommended content stored on the mobile phone. The mobile phone generates first recommendation information based on the recommended content corresponding to the first image resource and displays the first recommendation information in the search box.

[0439] In this implementation, the phone pre-generates recommendations for each image resource during its idle time. When the recommended information is needed, it searches for the recommended content corresponding to the selected image resource and generates the recommended information. This generation of recommended information does not require analyzing the image resource to obtain the recommended content; instead, the recommended content is directly read from the stored information, resulting in high efficiency and minimal latency.

[0440] In other embodiments, before displaying the search box (for example, in the scenario shown in FIG. 4C , after the mobile phone 100 receives a user click on the application icon 101), the mobile phone selects a first image resource from the gallery based on preset conditions (for example, the image resource must have been taken more than one month ago; for example, the recommended content must not be repeated within a week), performs data analysis on the first image resource, and generates recommended content corresponding to the first image resource. The mobile phone generates first recommendation information based on the recommended content corresponding to the first image resource and displays the first recommendation information in the search box.

[0441] In this implementation, when recommendation information needs to be displayed, an image resource is selected, semantic recognition is performed on the selected image resource to obtain recommended content and generate recommendation information. There is no need to save a large amount of recommended content, which can save storage space.

[0442] In the visual media search method provided by the embodiment of the present application, the electronic device displays recommended information to the user in the search box of the gallery. The recommended information is a sentence with natural semantics generated based on the first image resource in the gallery. The user can enter text information to search based on the content and format of the recommended information. For example, the user can use the recommended information as the text information in the search box to search. Since the recommended information is generated based on the first image resource in the gallery, the electronic device can search for the first image resource and all image resources with similar content to the first image resource in the gallery based on the text information, thereby improving the search hit rate. Moreover, since the recommended information is a sentence with natural semantics, it can more accurately express the purpose of the user's search, narrow the scope of the search results, and help users find the required pictures or videos more conveniently.

[0443] In some scenarios, users may not be able to quickly find the search box to search, but instead manually search for images or videos, which is inefficient. The visual media search method provided in the embodiments of the present application provides a method for quickly entering a search interface, guiding users to conveniently and quickly enter the search interface provided by the electronic device, and then the user can enter information in the search box of the search interface to automatically search for image resources.

[0444] Still taking the electronic device displaying the photo display interface as an example for introduction. Exemplarily, as shown in Figure 4D, the mobile phone 100 displays the photo display interface 401, and the photo display interface 401 includes a search box 402 and a photo display page 403. The search box 402 is used to trigger the display of the image resource search interface, and the photo display page 403 is used to display thumbnails of image resources (pictures or videos) in the gallery. The user may swipe up or swipe down on the photo display page 403. In one example, the photo display page 403 receives the user's upward swipe gesture, and in response to the user's upward swipe gesture, the mobile phone 100 displays the photo display interface 404, and the search box is no longer displayed on the photo display interface 404. Exemplarily, the photo display interface 404 includes a photo display page 405, and the photo display page 405 is used to display thumbnails of image resources (pictures or videos) in the gallery. In one implementation, the user continues to swipe upward on photo display page 405. In response to the user's swipe upward gesture on photo display page 405, the thumbnails of the image resources displayed on photo display page 405 scroll upward. For example, as shown in FIG4D , the thumbnails of the image resources on photo display page 403 are arranged from recent to recent according to the time of the image resources. Photo display page 403 displays image resources taken on November 9, 2023 and image resources taken on October 30, 2023. In response to the user's swipe upward gesture, the image resources taken on October 30, 2023 and image resources taken on September 17, 2023 are displayed on photo display page 405.

[0445] Continuing to refer to Figure 4D, in some embodiments, if it is detected that the upward sliding gesture in the photo display page 405 has not stopped, that is, the user continues to make the upward sliding gesture in the photo display page 405, the mobile phone 100 displays a prompt button 406, and the prompt button 406 displays a prompt message "Find photos, try the search function". The prompt message displayed on the prompt button 406 is used to prompt the user to click the prompt button 406 to trigger the display of the image resource search interface. Optionally, the mobile phone 100 displays the prompt button 406 in the photo display page 405. It should be noted that the continuous upward sliding gesture described in the embodiment of the present application refers to the mobile phone detecting the next upward sliding gesture within a certain period of time (such as 0.5 seconds) after detecting the end of an upward sliding gesture. If the next upward sliding gesture is not detected within a certain period of time (such as 0.5 seconds) after the end of an upward sliding gesture, it is determined that the upward sliding gesture has stopped.

[0446] The user can click prompt button 406 to trigger the display of the image resource search interface. For example, as shown in FIG4D , in response to the user clicking prompt button 406, mobile phone 100 displays search interface 201. Search interface 201 includes a search box 202, into which the user can enter text information. Upon detecting the user entering text information in search box 202, mobile phone 100 can search the image library for corresponding image resources based on the text information in search box 202.

[0447] In this way, when users are manually searching for photos or videos, they can easily enter the search interface by clicking a button according to the prompt information, enter text information in the search box of the search interface, and trigger the phone to automatically search for image resources in the gallery.

[0448] In one implementation, the search box on the search interface displays recommended information. This recommendation information is generated by the mobile phone based on the recommended content corresponding to the first image resource in the gallery and has natural semantics. For example, as shown in FIG4D , the search box 202 displays the recommendation message "Try searching for photos of a beach walk last August." Upon receiving the user input of "a beach walk last August" in the search box 202, the mobile phone 100 searches the gallery for image resources that match the semantics of "a beach walk last August."

[0449] 4E , in some embodiments, mobile phone 100 determines that an upward swipe gesture within photo display page 405 has stopped. In one example, if mobile phone 100 does not detect an upward swipe gesture within a certain period of time (e.g., 0.5 seconds), it determines that the upward swipe gesture within display page 405 has stopped. The thumbnails of image resources displayed within photo display page 405 stop scrolling upward. Exemplarily, as shown in FIG4E , mobile phone 100 displays photo display interface 407, which includes a search box 408 and a photo display page 409. Search box 408 is used to trigger the display of an image resource search interface, and photo display page 409 is used to display thumbnails of image resources (pictures or videos) in the gallery.

[0450] In one implementation, search box 408 displays recommended information. This recommended information is generated by the mobile phone based on the recommended content corresponding to the first image resource in the gallery and has natural semantics. For example, as shown in FIG4E , search box 408 displays the recommended information "Try searching for photos of a beach walk last August."

[0451] The user can click on search box 408 to trigger the display of an image resource search interface. For example, as shown in FIG4E , in response to the user clicking on search box 408, mobile phone 100 displays search interface 201. Search interface 201 includes search box 202, into which the user can enter text. Upon detecting the user entering text in search box 202, mobile phone 100 can search the image library for corresponding image resources based on the text in search box 202.

[0452] In this way, after the user stops manually searching for photos or videos, the phone displays a search box, prompting the user to enter text information to automatically search. The user conveniently enters the search interface by clicking the search box, and enters text information in the search box on the search interface to trigger the phone to automatically search for image resources in the gallery.

[0453] In one implementation, the search box on the search interface displays recommended information. This recommendation information is generated by the mobile phone based on the recommended content corresponding to the first image resource in the gallery and has natural semantics. For example, as shown in FIG4E , the search box 202 displays the recommendation message "Try searching for photos of a beach walk last August." Upon receiving the user input of "a beach walk last August" in the search box 202, the mobile phone 100 searches the gallery for image resources that match the semantics of "a beach walk last August."

[0454] It should be noted that the embodiment of the present application is introduced by taking the example of a mobile phone detecting a user's upward sliding gesture on the photo display interface, and this example should not constitute a limitation on the applicable scenarios of the embodiment of the present application. In other embodiments, the applicable scenarios of the embodiment of the present application also include: the mobile phone detects an upward sliding gesture on other pages in the gallery for displaying thumbnails of image resources to users, or detects a downward sliding gesture on other pages for displaying thumbnails of image resources to users. The specific implementation methods of these scenarios can refer to the implementation methods of detecting an upward sliding gesture on the photo display interface shown in Figures 4D and 4E. The implementation methods are similar and will not be shown one by one in the embodiment of the present application.

[0455] The visual media search method provided by the embodiment of the present application is that when the user swipes up or down to search for pictures or videos, the mobile phone displays a prompt button and prompt information, prompting the user to trigger the mobile phone to display a search interface, so that the mobile phone automatically searches for image resources based on the text information entered by the user in the search box. After the user finishes swiping up or down to search for pictures or videos, the mobile phone displays a search box and recommended information, prompting the user to trigger the mobile phone to display a search interface, so that the mobile phone automatically searches for image resources based on the text information entered by the user in the search box. The mobile phone prompts the user in various ways to more conveniently use the mobile phone to automatically search for image resources, thereby improving the convenience of users in finding image resources in the gallery.

[0456] The mobile phone receives text information input by the user in the search box on the gallery user interface, and searches for image resources corresponding to the text information according to the text information. These image resources may include pictures and videos.

[0457] In one implementation, a text analysis algorithm is used to perform semantic understanding on the text information input by the user to obtain the constituent elements corresponding to the text information, including at least one of time, place, person, subject type and event.

[0458] For images, the time and location of the photo can be obtained based on the image information stored on the phone. Semantic analysis algorithms can also be used to analyze the image semantics, extracting one or more of the following: person, subject type, and event. This allows us to identify the image's constituent elements.

[0459] If it is determined that the constituent elements included in the text information are consistent with the constituent elements included in the picture, then it is determined that the picture is the image resource corresponding to the text information.

[0460] As can be understood, a video is composed of a number of image frames. In one implementation, if it is determined that a video includes N image frames corresponding to text information, then the video is determined to be an image resource corresponding to the text information. Where N is a preset value, for example, N=1, N=24, etc.

[0461] The mobile phone determines the pictures and videos corresponding to the text information based on the text information and can display the search results to the user. For example, thumbnails of these pictures and videos can be displayed in the user interface of the gallery.

[0462] As shown in FIG4F , mobile phone 100 displays a search interface 201, which includes a search box 202. A user enters the text "last August beach walk" in search box 202 and then clicks search box 202. Mobile phone 100 receives the user's input of "last August beach walk" in search box 202 and searches the gallery for image resources that are semantically consistent with "last August beach walk." As shown in FIG4F , mobile phone 100 displays a search results interface 501, which displays image resources found in the gallery based on the text in search box 202. For example, 54 image resources in the gallery that are semantically consistent with "last August beach walk" were found. Alternatively, in one example, search results interface 501 includes pages 502, 503, and 504. Page 502 displays images and videos corresponding to the user's input text, page 503 displays images corresponding to the user's input text, and page 504 displays videos corresponding to the user's input text.

[0463] In some embodiments, the video corresponding to the text information includes image frames corresponding to the text information and image frames not corresponding to the text information. For example, as shown in FIG4G , video A is a video corresponding to “Walking on the beach last August” searched by mobile phone 100. Video A includes many image frames, among which the image frames from 00:00 to 02:17 do not correspond to “Walking on the beach last August”. For example, the image frames from 00:00 to 02:17 are image frames corresponding to “Playing volleyball on the beach last August”; the image frames from 02:18 to 06:21 are image frames corresponding to “Walking on the beach last August”; and the image frames from 06:22 to 08:32 do not correspond to “Walking on the beach last August”. For example, the image frames from 06:22 to 08:32 are image frames corresponding to “Running on the beach last August”.

[0464] In one implementation, the video thumbnail image displayed by the mobile phone in the search result interface is the first image frame in the video, and the first image frame is one of the image frames corresponding to the text information.

[0465] In an example, the first image frame is any one of the image frames corresponding to the text information.

[0466] In another example, the first image frame is the first image frame corresponding to the text message. In one implementation, the mobile phone determines the first image frame corresponding to the text message frame by frame according to the chronological order of the image frames in the video. When the first image frame corresponding to the text message is found, the first image frame corresponding to the text message is determined as the first image frame. Using frame-by-frame determination, the first image frame corresponding to the text message can be accurately determined.

[0467] In another example, each video is divided into one or more video segments according to a preset condition; for example, the preset condition is that a video segment is divided every 0.5 seconds according to the playback time, or a video segment is divided every M (for example, M=12) frames. Each video segment includes a representative frame, which can be used to represent the video segment. For example, the representative frame is the first frame in a video segment, or the representative frame is the last frame in a video segment, or the representative frame is the frame with the highest pixel in a video segment, or the representative frame is the optimal frame in a video segment, or the representative frame is any frame in a video segment, etc. If the representative frame of a video segment is an image frame corresponding to text information, then the video segment is determined to be a video segment corresponding to the text information. In one implementation, the first image frame is an image frame in the first video segment corresponding to the text information (for example, the image frame is the representative frame in this video segment). For example, the 276th to 762nd video segments in video A are video segments corresponding to text information. The representative frame of the 276th video segment in the video A is the first image frame. For example, the representative frame of the 276th video segment can be the first frame of the 276th video segment, or the last frame of the 276th video segment, or the frame with the highest pixel in the 276th video segment, etc.

[0468] For example, as shown in FIG4F , the thumbnail 5021 displayed in the page 502 is a thumbnail of the video A corresponding to “A walk on the beach last August”, and the image of the thumbnail 5021 is the first image frame in the video A.

[0469] For example, as shown in FIG4F , the thumbnail 5041 displayed in the page 504 is the thumbnail of the video A corresponding to “A walk on the beach last August”, and the image of the thumbnail 5041 is the first image frame in the video A.

[0470] In some embodiments, the mobile phone displays the position of the first image frame in the video on the search results interface. For example, the position of the first image frame in the video can be indicated by the playback time corresponding to the first image frame and the total playback time of the video. This allows the user to easily find the content they are searching for without having to manually search for the desired image frame in the video.

[0471] For example, the playback time of the first image frame is 02 minutes and 18 seconds, and the total playback time of Video A is 08 minutes and 32 seconds. As shown in Figure 4F, "02:18 / 08:32" is superimposed on the image of thumbnail 5021, indicating that the image of thumbnail 5021 is a frame at 02 minutes and 18 seconds in the total playback time of 08 minutes and 32 seconds. The corresponding position of thumbnail 5041 (next to thumbnail 5041) displays "02:18 / 08:32", indicating that the image of thumbnail 5041 is a frame at 02 minutes and 18 seconds in the total playback time of 08 minutes and 32 seconds.

[0472] In the visual media search method provided by the embodiments of the present application, the search results interface displays the position of the first image frame in the video, allowing the user to easily locate the position in the video corresponding to the text information entered by the user. For example, in FIG4F , the text information entered by the user is "Walking on the beach last August," and the playback time of the first image frame corresponding to "Walking on the beach last August" is 02 minutes and 18 seconds. The thumbnail of video A displayed in the search results interface displays this time point (02 minutes and 18 seconds), and the image of the thumbnail of video A displayed in the search results interface is the image frame corresponding to 02 minutes and 18 seconds (the first image frame). In another example, as shown in FIG4H , the text information entered by the user in input box 202 is "Running on the beach last August," and the playback time of the first image frame corresponding to "Running on the beach last August" is 06 minutes and 21 seconds. In response to receiving the user input "Running on the beach last August," mobile phone 100 displays search results interface 505, which is used to display image resources found in the image library based on the text information in search box 202. For example, a total of 68 image resources with semantic consistency with "Running on the beach last August" were found in the image library. Optionally, in one example, search results interface 505 includes pages 506, 507, and 508; page 506 is used to display images and videos corresponding to the text information entered by the user, page 507 is used to display images corresponding to the text information entered by the user, and page 508 is used to display videos corresponding to the text information entered by the user. Thumbnail 5061 displayed on page 506 is a thumbnail of video A corresponding to "Running on the beach last August," and the image in thumbnail 5061 is the first image frame in video A determined based on "Running on the beach last August." Thumbnail 5081 displayed on page 508 is a thumbnail of video A corresponding to "Running on the beach last August," and the image in thumbnail 5081 is the first image frame in video A determined based on "Running on the beach last August." For example, the first image frame corresponding to "Running on the beach last August" in video A is played at the 6th minute and 21st second mark, and the total playback duration of video A is 8 minutes and 32 seconds. As shown in FIG4H , "06:21 / 08:32" is superimposed on thumbnail 5061, indicating that thumbnail 5061 is a frame at the 6th minute and 21st second mark of the video with a total playback duration of 8 minutes and 32 seconds. The corresponding position of thumbnail 5081 (next to thumbnail 5081) displays "06:21 / 08:32," indicating that thumbnail 5081 is a frame at the 6th minute and 21st second mark of the video with a total playback duration of 8 minutes and 32 seconds.

[0473] In some embodiments, in response to a user clicking on any video thumbnail in the search results interface, the mobile phone plays the video corresponding to the video thumbnail. In one implementation, the mobile phone starts playing the video corresponding to the video thumbnail from the position of the image frame (first image frame) corresponding to the image of the video thumbnail.

[0474] For example, as shown in FIG4J , mobile phone 100 displays search results interface 501, which displays image resources in the gallery corresponding to the user-entered text message "Walking on the beach last August." Search results interface 501 includes pages 502, 503, and 504. Page 502 includes a thumbnail 5021 of video A. Thumbnail 5021 is the first frame of video A. The playback time of the first frame is 02 minutes and 18 seconds, and the total playback time of video A is 08 minutes and 32 seconds.

[0475] In response to a user clicking on thumbnail 5021, mobile phone 100 plays video A and displays playback interface 601 of video A. In one implementation, the mobile phone starts playing video A from the first image frame. For example, as shown in FIG4J , in response to a user clicking on thumbnail 5021, mobile phone 100 displays playback interface 601 of video A. Playback interface 601 includes a progress bar 602. For example, progress bar 602 is positioned directly at the 02 minute 18 second mark, and mobile phone 100 starts playing video A at the 02 minute 18 second mark.

[0476] In this method, the mobile phone starts playing the video from the image frame (first image frame) corresponding to the text information input by the user, and can directly locate the content that the user needs to find, which is convenient and fast, and improves the user experience.

[0477] In some embodiments, a video includes multiple video segments corresponding to text information input by the user. For example, as shown in FIG4K , video A includes two video segments corresponding to the text information "Walking on the beach last August." The first video segment includes image frames from 02 minutes 18 seconds to 06 minutes 21 seconds, and the second video segment includes image frames from 07 minutes 03 seconds to 08 minutes 32 seconds.

[0478] Accordingly, the start position of each video segment corresponding to the text information entered by the user is marked on the progress bar 602 in the playback interface 601. For example, as shown in FIG4K , the progress bar 602 includes a marking point 603 and a marking point 604. Marking point 603 indicates the start time of the first video segment in Video A corresponding to "Walking on the Beach Last August," and marking point 604 indicates the start time of the second video segment in Video A corresponding to "Walking on the Beach Last August." This allows the user to clearly see the start time of all video segments in Video A corresponding to the text information entered by the user.

[0479] In one implementation, the mobile phone starts playing video A from the start moment of the first video segment corresponding to the text information input by the user in video A. For example, as shown in FIG4K , the mobile phone 100 starts playing video A from the position indicated by the marked point 603 .

[0480] In one implementation, when the mobile phone receives a user click operation on any marked point on the progress bar 602, the mobile phone 100 starts playing video A from the position indicated by the marked point. For example, referring to FIG4K, if the mobile phone 100 receives a user click operation on the marked point 604, the mobile phone 100 starts playing video A from the position indicated by the marked point 604.

[0481] In this method, each video segment corresponding to the text information entered by the user can be marked in the video playback progress bar. This allows users to easily know the starting time of all video segments corresponding to the text information entered by the user. Furthermore, users can click any marked point to have the phone start playing the video from the corresponding position of the marked point, allowing users to easily find the video segment they are looking for.

[0482] The following describes the implementation process of the visual media search method provided in the embodiment of the present application in conjunction with the interaction diagram corresponding to the above-mentioned interface changes.

[0483] The visual media search method provided by the present application is described in detail below with reference to Figure 5a. As shown in Figure 5a, the visual media search method includes an index building phase and a search phase.

[0484] During the index building phase, in some embodiments, an index may be built for each video frame of the video.

[0485] In some embodiments, if this solution is implemented on a terminal device, an index is built for each video frame. During the search phase, frame-by-frame matching is required, which increases search latency and reduces user experience. Alternatively, in some embodiments, the video can be segmented and a video index can be built per segment.

[0486] For example, as shown in FIG5a, the video is first deframed to obtain multiple video frames of the video; then, a video segmentation algorithm is used to segment the video based on the multiple video frames of the video to obtain multiple video segments such as video segment 1 and video segment 2; then, a representative frame can be selected from the multiple video frames in the video segment, and the scores of the multiple video frames in the video segment are evaluated respectively, and the video frame with the highest score is determined as the representative frame of each video segment; the representative frame is input into the image encoder of the CLIP model, and the visual semantic vector of the representative frame is output; an index is constructed based on the visual semantic vector corresponding to the representative frame, and an inverted index library is obtained based on the constructed index.

[0487] It should be noted that the above-mentioned method of selecting representative frames is only an example. The starting frame, middle frame, end frame or a random video frame of the video segment can also be selected as the representative frame. The specific implementation method can be found in the description of the embodiment below. In the search stage, the user enters the search text in the search interface; the search text is input into the text encoder of the CLIP model, and the text semantic vector corresponding to the search text is output; based on the text semantic vector corresponding to the search text, vector recall is performed from the inverted index library; the search text is sent to the natural language understanding module; the natural language understanding module performs entity recognition on the search text to obtain entities in the search text; based on the entities in the search text, entity recall is performed from the inverted index library; based on the vector recall and entity recall, vector recall results and entity recall results are returned from the inverted index library; the vector recall results and search results are sorted to obtain search results; and the search results are displayed on the search interface.

[0488] In some embodiments, the image encoder of the CLIP model and the text encoder of the CLIP model can be trained based on contrastive learning.

[0489] FIG5b below still takes the electronic device being a mobile phone 300 as an example, and combines the gallery service module, search module, multimodal understanding module and natural language understanding module displayed in the system library in FIG2b to explain in detail the visual media search method provided in the embodiment of the present application.

[0490] As shown in FIG5 b , the visual media search method provided in the embodiment of the present application can be divided into the following two stages: an index building stage and a search stage.

[0491] First, the steps included in the index building phase are described in detail with reference to FIG5 b .

[0492] S501: The gallery service module 510 receives a user's operation of adding or modifying a video.

[0493] Adding a video refers to a user storing a video in the Gallery app. For example, a user adding a video may include shooting a video, downloading a video, or recording the screen of the phone 300 using the phone 300. Modifying a video refers to a user modifying a video already stored in the Gallery app. For example, a user modifying a video may include cropping, splicing, adding special effects, or adding subtitles.

[0494] In some embodiments, the gallery service module 510 receiving a user's new operation or modification operation on a video includes: the gallery service module 510 receiving a user's new operation or modification operation on an attribute tag of a video.

[0495] For example, a video's attribute tags may include the video's capture location (e.g., the location where the video was shot, the source from which the video was downloaded, etc.), the video's capture time (e.g., the time it was shot, downloaded, or the duration of the screen recording), names of people added or modified by the user for the video, or classification tags, events, and so on. Classification tags can be used to indicate the type of object depicted in the video. For example, classification tags can include people, plants, animals, buildings, or natural scenery. Events can be used to indicate what the object depicted in the video does. For example, events can include games, sports, and so on.

[0496] In some embodiments, the classification label may be manually configured by a user, or may be obtained by the electronic device automatically classifying the video.

[0497] S502: The gallery service module 510 stores the video.

[0498] In response to a user's operation to add or modify a video, the gallery service module 510 stores the video on the mobile phone 300. For example, the video can be stored in a local file on the mobile phone 300, and the user can access the video through various channels, such as the local folder on the mobile phone 300 or the gallery application. For another example, with user authorization, the video can be stored in the cloud for backup, reducing the memory pressure on the mobile phone 300.

[0499] In some embodiments, in response to a user's operation of adding or modifying attribute tags of a video, the gallery service module 510 may store the attribute tags of the video.

[0500] S503: The gallery service module 510 calls the multimodal understanding module 530 to determine a representative frame of the video.

[0501] It's understandable that a video typically consists of multiple frames. If we determine the corresponding visual semantic vector for each frame and store it as an index, each video would correspond to a large number of indexes, wasting significant computing resources and storage space. Furthermore, these large indexes may contain noise, which can lead to time-consuming search and matching, potentially affecting video search results.

[0502] Therefore, in the embodiment of the present application, the multimodal understanding module 530 first segments the video and determines the corresponding representative frame from each video segment. This allows the subsequent steps to determine the corresponding visual semantic vector for the representative frame and store it as an index, greatly reducing the number of indexes corresponding to the video. In this way, the representative frames corresponding to multiple video segments can more completely represent the video semantics of the video, while also saving costs.

[0503] In some embodiments, the gallery service module 510 may call the computer vision service provided by the multimodal understanding module 530 to determine the representative frame of the video. The computer vision service refers to the multimodal understanding module 530 performing video semantic understanding on the video, and then the multimodal understanding module 530 determines the representative frame of the video.

[0504] Video semantic understanding refers to enabling mobile phones to understand the meaning of the content displayed in the video, such as understanding the type, quantity, location, and relationship between objects in the video.

[0505] In some embodiments, the computer vision service provides services such as video frame decomposition, video segmentation using a video segmentation algorithm, and determination of representative frames for each video segment. The process of determining the representative frame of a video can be broken down into the following steps: 1-3:

[0506] Step 1: The multimodal understanding module 530 uses computer vision services to perform frame decomposition processing on the video.

[0507] It is understood that a video is composed of multiple video frames, each of which is a still image in the video. In other words, each video frame can be regarded as an image. Frame decomposition processing refers to the multimodal understanding module 530 using computer vision services to decompose the video into individual video frames.

[0508] In some embodiments, during the video de-framing process performed by the multimodal understanding module 530, each resulting video frame may be labeled with a frame identifier and a time point. A frame identifier uniquely identifies a video frame, and different video frames can be distinguished based on the frame identifier. A time point refers to the time at which a video frame appears in the video.

[0509] In some embodiments, during the process of de-framing a video using computer vision services, the multimodal understanding module 530 can identify a classification label for each video frame. The classification label can be used to indicate the type of object displayed in the video frame. For example, the classification label can be a person, plant, or animal displayed in the video frame. Another example is a building, natural scenery, or the like displayed in the video frame.

[0510] It should be noted that there is no limit to the number of category labels for video frames. For example, as shown in FIG3 a , the category labels of the thumbnail of “Video A” included in the first search result display interface 330 may include “Boy”, “Stars”, “Sky”, and “Tree”.

[0511] Step 2: The multimodal understanding module 530 uses computer vision services to segment the video.

[0512] Segmentation processing refers to dividing a video into multiple video segments using the video segmentation algorithm provided by the computer vision service.

[0513] It is understandable that when a video is played, the content displayed changes continuously as the video frames are played in sequence, but the degree of change is different. Therefore, similar video frames (with a small degree of change) can be classified into one video segment.

[0514] In some embodiments, the degree of change between a video frame and its adjacent frames can be measured based on the classification labels and image parameters of the video frames, thereby segmenting the video. For example, the image parameters of the video frames may include jitter, clarity, pixel value, etc., which are not limited in this application.

[0515] Video frame jitter refers to the phenomenon of jitter or shaking of the content displayed in the video frame during video playback. For example, when a user holds mobile phone 300 to shoot a video and then moves phone 300 to capture another scene, noticeable jitter may occur. Video frame clarity refers to the clarity of detailed textures and their boundaries within the video frame. The number of pixels in a video frame can represent the brightness of the video frame.

[0516] It can be understood that, compared with its adjacent video frames, the value corresponding to the jitter of a video frame, the degree of change of other image parameters and the degree of change of the classification label are positively correlated with the degree of change of the content displayed in the video. That is, the larger or smaller the value corresponding to the jitter, the more obvious the degree of change of the image parameters and the degree of change of the classification label, and the more obvious the degree of change of the content displayed in the video.

[0517] In some embodiments, the video segmentation algorithm can be expressed using the following formula 1: based on the classification label and image parameters of the video frame, the segmentation score of the video frame is calculated in combination with the following formula 1, and then the segmentation score of the video frame is used to determine whether it is determined as the starting frame of a video segment or the ending frame of a video segment. Formula 1 is as follows: y = α × frame A +β×frame B +γ×frame Y +δ×frame T

[0518] Among them, y represents the segmentation score of the video frame, frame A Indicates the jitter score of the video frame, frame B Indicates the clarity change score of the video frame, frame Y Indicates the label change score of the video frame, frame T Indicates the pixel change fraction of the video frame, α, β, γ and δ represent the frame A 、frame B 、frame Y and frame T In some embodiments, α, β, γ, and δ may be manually preset values ​​based on the degree of influence of jitter, definition change, label change, and pixel change of the video frame on the segmentation score of the video frame.

[0519] It should be noted that the above process of determining the segmentation score of a video frame based on the classification label and multiple image parameters is only an example. The segmentation score of a video frame can also be determined based on the classification label and one or more of the multiple image parameters, and this application does not limit this.

[0520] In some embodiments, the jitter levels of multiple video frames in a video can be detected based on methods such as optical flow based on image displacement, feature point matching, and image grayscale distribution features. The jitter score of each video frame can then be determined based on a pre-set correspondence between the jitter range and the jitter score.

[0521] In some embodiments, a clarity detection tool can be used to determine the clarity of multiple video frames in a video, and then the clarity of the video frame is compared with the clarity of adjacent video frames to determine the clarity change value of the video frame. Subsequently, based on the correspondence between a pre-set clarity change value range and the clarity change score, the clarity change score of the video frame is determined. Exemplarily, the video quality detection tool can be open source software such as FFmpeg or Video Quality Measurement Tool.

[0522] In some embodiments, the classification labels of a video frame can be compared with the classification labels of adjacent video frames, and the comprehensive changes in the labels of the video frames can be determined based on the changes in the number of classification labels of the video frames and the changes in the content of the classification labels. For example, the union and intersection of the classification labels of a video frame and the classification labels of its adjacent video frames can be calculated. The change in the number of classification labels in the union can reflect the change in the number of classification labels of the video frame, and the change in the number of classification labels in the intersection can reflect the change in the content of the classification labels. When there is a change in the number of classification labels in the union or the number of classification labels in the intersection, the label change score of the video frame can be determined based on the correspondence between the pre-set range of the change in the number of classification labels and the label change score.

[0523] In some embodiments, a pixel detection tool can be used to determine the pixels of multiple video frames in a video, and then the pixels of the video frame are compared with the pixels of adjacent video frames to determine the pixel change value of the video frame. The pixel change score of the video frame is then determined based on the correspondence between a pre-defined pixel change value range and the pixel change score. Exemplarily, the pixel detection tool can be a plug-in such as PixelStick, MeasureIt, or Guides.

[0524] In some embodiments, the degree of change in the content displayed by two adjacent video frames in a video may be positively correlated with the size of the segmentation score of the video frame, that is, the larger the segmentation score of the video frame, the more obvious the degree of change in the content displayed by it compared with the adjacent video frames.

[0525] For example, a video frame can be compared with its previous video frame, and a segmentation score threshold is pre-set, and the video frame whose segmentation score exceeds the segmentation score threshold is determined as the starting frame of a video segment. Combined with the above formula 1, the higher the jitter of the video frame, the higher the frame jitter. A The higher the corresponding value, the greater the change in clarity of the video frame compared to the previous video frame. B The higher the corresponding value, the greater the label change of this video frame compared with the previous video frame. Y The higher the corresponding value, the greater the label change of this video frame compared with the previous video frame. Y The higher the corresponding value, the greater the pixel change between the video frame and the previous video frame. T The higher the corresponding value.

[0526] As another example, a video frame may be compared with its next video frame, and a segmentation score threshold may be preset, and a video frame whose segmentation score exceeds the segmentation score threshold may be determined as an end frame of a video segment.

[0527] In some embodiments, the degree of change in the content displayed by two adjacent video frames in a video may also be negatively correlated with the size of the segmentation score of the video frame, that is, the smaller the segmentation score of the video frame, the more obvious the degree of change in the content displayed by it compared with the adjacent video frames, that is, the lower the segmentation score of the video frame, the greater the possibility of determining it as the starting frame of a video segment or the ending frame of a video segment.

[0528] For example, a video frame can be compared with its previous video frame, and a segmentation score threshold is pre-set, and the video frame with a segmentation score lower than the segmentation score threshold is determined as the starting frame of a video segment. Combined with the above formula 1, the higher the jitter of the video frame, the higher the frame jitter. A The lower the corresponding value, the greater the change in clarity between the video frame and the previous video frame. B The lower the corresponding value, the greater the label change of this video frame compared with the previous video frame. Y The lower the corresponding value, the greater the label change of this video frame compared with the previous video frame. Y The lower the corresponding value, the greater the pixel change between the video frame and the previous video frame. T The corresponding value is lower.

[0529] Step 3: The multimodal understanding module 530 uses computer vision services to determine the representative frame corresponding to each video segment.

[0530] After obtaining multiple video segments, the multimodal understanding module 530 determines, from each video segment, a representative frame that can represent the video semantics of the video segment.

[0531] In some embodiments, the representative frame may be a start frame, an end frame, an intermediate frame, or a random frame in a video segment.

[0532] For example, assuming that a video segment contains 99 video frames, the first video frame (starting frame), the 99th video frame (ending frame), or the 50th video frame (middle frame) of the 99 video frames can be used as the representative frame of the video segment according to the time sequence, or a video frame (random frame) can be randomly selected from the 99 video frames as the representative frame of the video segment.

[0533] In some embodiments, an optimal frame may be determined from a video segment as a representative frame based on preset rules related to image parameters and classification labels of the video frames.

[0534] For example, a video frame score can be calculated based on its image parameters and classification labels. The more labels a video frame has, the higher its score. The lower the jitter of a video frame, the higher its score. The higher the clarity of a video frame, the higher its score. The lower the pixel change between a video frame and the previous frame, the higher its score. Finally, the video frame with the highest score (i.e., the optimal frame) in a video segment can be determined as the representative frame.

[0535] Next, the process of determining the representative frame in the video will be described in detail with reference to FIG6 .

[0536] As shown in FIG6 , the video includes 120 video frames. A video frame is compared with the previous video frame, and the segmentation scores corresponding to the 120 video frames are calculated using the above formula 1. The video frame whose segmentation score exceeds the segmentation score threshold is used as the starting frame of a video segment.

[0537] In some embodiments, the 120 video frames are divided into three video segments: video segment 1, video segment 2, and video segment 3. The middle frame of each video segment is determined as its corresponding representative frame. Video segment 1 includes 40 video frames, video segment 2 includes 51 video frames, and video segment 3 includes 29 video frames. That is, the 20th video frame in video segment 1 is used as the representative frame corresponding to video segment 1, the 26th video frame in video segment 2 is used as the representative frame corresponding to video segment 2, and the 15th video frame in video segment 3 is used as the representative frame corresponding to video segment 3.

[0538] It should be noted that the number of video frames in video segment 1 is 40, which is an even number. The 20th video frame and the 21st video frame are both intermediate frames. Any one of the intermediate frames can be determined as the representative frame corresponding to video segment 1. This application does not limit this.

[0539] S504 : The multimodal understanding module 530 returns the representative frame of the video and related information to the gallery service module 510 .

[0540] The relevant information about the representative frame includes, but is not limited to, the time points corresponding to the start and end frames of the video segment corresponding to the representative frame, the time points corresponding to the representative frame, and the classification label of the video segment. The time points corresponding to the start, end, and representative frames of the video segment facilitate jumping to the corresponding video frames when subsequently displaying video search results to the user, improving the user experience. The classification label of the video segment facilitates displaying more accurate video search results during subsequent video searches.

[0541] In some embodiments, the classification label of a video segment can be used to indicate the type of object displayed in the video segment. For example, the classification label of the video segment can be obtained by taking the union of the classification labels corresponding to the multiple video frames included in the video segment.

[0542] In addition, the relevant information of the representative frame may also include names and character relationships corresponding to the video segments. In some embodiments, after determining the representative frame, the names and character relationships corresponding to the characters in the representative frame may be identified based on pre-set names and character relationships.

[0543] S505: The gallery service module 510 stores representative frames of the video and related information.

[0544] After receiving the representative frames and related information returned by the multimodal understanding module 530, the gallery service module 510 may store them. In some embodiments, the gallery service module 510 is configured with a database, and the gallery service module 510 may store the representative frames of the video and related information in the database.

[0545] S506: The gallery service module 510 calls the multimodal understanding module 530 to determine the visual semantic vector corresponding to the representative frame.

[0546] In some embodiments, the image library service module 510 uses the computer vision service provided by the multimodal understanding module 530 to determine the visual semantic vectors corresponding to the representative frames. A video includes multiple video segments, each of which corresponds to a representative frame. Therefore, a video corresponds to multiple visual semantic vectors. The image library service module 510 uses the computer vision service provided by the multimodal understanding module 530 to perform video semantic understanding on the video and further determine the visual semantic vectors corresponding to the multiple representative frames in the video.

[0547] In some embodiments, the computer vision service may provide a CLIP model, wherein a representative frame may be input into an image encoder of the CLIP model to obtain a visual semantic vector corresponding to the representative frame.

[0548] In some embodiments, the image encoder of the CLIP model can be trained using image training samples, which may include picture training samples and video frame training samples. The training process of the image encoder of the CLIP model can be seen in the following embodiments.

[0549] S507: The multimodal understanding module 530 returns the visual semantic vector corresponding to the frame to the gallery service module 510.

[0550] S508: The gallery service module 510 stores the visual semantic vector corresponding to the representative frame.

[0551] The image library service module 510 receives the visual semantic vector corresponding to the representative frame returned by the multimodal understanding module 530 and can store it. For example, the image library service module 510 is configured with a database, and the image library service module 510 can store the visual semantic vector corresponding to the representative frame in the database.

[0552] S509 : The gallery service module 510 sends the representative frame of the video and its related information as well as the visual semantic vector of the representative frame to the search module 520 .

[0553] S510: The search module 520 constructs an index corresponding to the video.

[0554] An index is the index corresponding to a video segment in a video. If a video includes multiple video segments, then a video corresponds to multiple indexes, and different indexes correspond to different video segments in the video.

[0555] The search module 520 can construct an index of the video segment corresponding to the representative frame by combining the visual semantic vector of the representative frame, the relevant information of the representative frame, and the attribute label of the video segment. As in the example above, the index of the video segment can include the visual semantic vector of the representative frame of the video segment, the time point corresponding to the starting frame, the time point corresponding to the ending frame, and the time point corresponding to the representative frame in the video segment, the classification label of the video segment, and the attribute label of the video segment (such as the video acquisition time, video acquisition location, and video storage path).

[0556] It should be noted that the content included in the index of the video segment is only an example, and the index of the video segment may include the visual semantic vector representing the frame, as well as related information representing the frame and any information in the attribute tag of the video segment. This application does not limit this.

[0557] Exemplarily, the search module 520 may include an index library, and the search module 520 may store the index corresponding to the video in the index library, so that subsequent search and matching can be performed based on the index library.

[0558] It is understandable that, in actual application, the number of videos stored in mobile phone 300 may be very large, and one video may correspond to multiple indexes, indicating that the index library may contain a large number of indexes. In some embodiments, in order to improve search efficiency, an inverted index may be used to search.

[0559] For example, after the search module 520 constructs an index corresponding to a video, it performs vector clustering on the visual semantic vectors included in the index in the index library to obtain multiple clusters. Specifically, the vector space corresponding to all indexes is divided into multiple vector regions, each of which includes a cluster. Each vector region includes multiple indexes with high vector similarity, and each vector region can be represented by a cluster center. For example, clustering can be performed using methods such as K-means or hierarchical clustering, which is not limited in this application.

[0560] In this way, when the mobile phone 300 performs search matching based on the text semantic vector of the search text, it can first match with the cluster center point, and then match with the visual semantic vector in the vector area to which the determined cluster center point belongs. There is no need to match with all indexes in the index library. This saves computing resources and greatly reduces the time consumed by search matching, avoids search delays, improves video search efficiency, and further enhances the user experience.

[0561] In some embodiments, the search module 520 may store the received representative frames and related information thereof as well as the visual semantic vectors of the representative frames to facilitate searching when building an index, thereby improving the speed of building the index.

[0562] For example, the visual semantic vector of the representative frame corresponding to a video segment and related information of the representative frame can be stored in a document. Taking video segment 1 as an example, the visual semantic vector of the representative frame corresponding to video segment 1 and related information of the representative frame can be stored in sub-document 1. The information stored in sub-document 1 can be shown in Table 3 below:

[0563] Table 3

[0564] As shown in Table 3, segments.media-vector represents the visual semantic vector representing the frame in video segment 1. In practical applications, vectors are usually composed of arrays. In an embodiment of the present application, the visual semantic vector is processed in hexadecimal sequence to obtain a visual semantic vector in the form of [bd de d1 b4 3c 8c 9c], and the hexadecimal floating point number can enable the mobile phone 300 to store with less storage space, which can save storage space. segments.startTime represents the starting time of video segment 1 0ms (that is, the time point corresponding to the starting frame in video segment 1). segments.endTime represents the ending time of video segment 1 91666ms (that is, the time point corresponding to the ending frame in video segment 1). segments.startFrame and segments.endFrame indicate that there are a total of 2750 frames of video frames in video segment 1 from the start to the end. segments.tag-name represents the classification label of video segment 1, including "character", "scenery" and "architecture".

[0565] In some embodiments, the attribute tag of video segment 1 may also be stored in sub-document 1. The attribute tag of video segment 1 is the attribute tag of the video stored in step S502. For example, if video segment 1 is a video segment in a video shot by a user, the attribute tag of video segment 1 may include the shooting time and location of video segment 1, and the storage path of mobile phone 300.

[0566] In some embodiments, the attribute tags of video segment 1 may also be stored in document 1. It is understood that the attribute tags of a video are fixed, meaning that the attribute tags corresponding to multiple video segments in a video are the same. Therefore, the attribute tags of a video may be stored in document 1, and the attribute tags of each video segment of the video may be retrieved from document 1, thereby reducing content load on mobile phone 300 and reducing storage costs.

[0567] For example, video D includes the above-mentioned video segment 1, as well as video segments 2 and 3. The visual semantic vectors and related information of the representative frames corresponding to video segment 1 are stored in sub-document 1, the visual semantic vectors and related information of the representative frames corresponding to video segment 2 are stored in sub-document 2, and the visual semantic vectors and related information of the representative frames corresponding to video segment 1 are stored in sub-document 3. The information stored in document 1 can be shown in Table 4 below:

[0568] Table 4

[0569] As shown in Table 4, file_path represents the storage path of video D on mobile phone 300. shooting-time represents the shooting time of video D. location represents the shooting location of video D. segments indicates the subdocuments corresponding to the multiple video segments included in video D. Based on Tables 3 and 4 above, the attribute tags of video segment 1 can be obtained from document 1, and information related to the representative frames of video segment 1 can be obtained from subdocument 1. Similarly, the attribute tags of video segments 2 and 3 can also be obtained from document 1.

[0570] It should be noted that, since video semantic understanding may require a large amount of computing resources, in order to avoid causing user experience lags, etc., steps S503-S510 can be executed when the mobile phone 300 is in the charging and screen-off state.

[0571] The visual media search method provided by the present application is based on searching and matching the search text input by the user with the index corresponding to the video to obtain search results. The index corresponding to the video is the index of a video segment in the video, and the index of the video segment at least includes the visual semantic vector representing the frame in the video segment. The visual semantic vector representing the frame can indicate the meaning expressed by the picture content displayed by the representative frame, which can represent the video semantics of the video segment. Therefore, the visual semantic vector representing the frame is fully associated with the video semantics of the video segment in the video. Search matching based on the visual semantic vector representing the frame can realize the fusion interaction of the search text and the picture content displayed by the representative frame, thereby improving the accuracy of the video search results and enhancing the user experience.

[0572] Next, the steps included in the search phase of the visual media search method provided by the present application will be described in detail with reference to FIG5 b .

[0573] S511: The gallery service module 510 receives a user input operation for a search text.

[0574] The user can enter search text in the search interface provided by the mobile phone 300. For example, the user can enter search text in the search box 321 included in the album display interface 320 as shown in Figure 3a. For another example, the user can enter search text in the search box 421 included in the negative one screen interface 420 as shown in Figure 4a.

[0575] The search text is a text describing the characteristics of the video that the user wants. For example, the search text may include the time the video was acquired, the location where the video was acquired, and the content of the video. For example, the search text may be "scenery taken last week", which is not limited in this application.

[0576] S512 : The gallery service module 510 sends the search text to the search module 520 .

[0577] S513: The search module 520 calls the multimodal understanding module 530 to determine the text semantic vector corresponding to the search text.

[0578] The search module 520 calls the multimodal understanding module 530 to perform text semantic understanding on the search text to obtain a text semantic vector corresponding to the search text.

[0579] Text semantic understanding refers to enabling mobile phones to understand the meaning expressed in text, and is a key technology in natural language processing (NLP) technology.

[0580] In some embodiments, the multimodal understanding module 530 provides a CLIP model, and the search text can be input into the text encoder of the CLIP model to obtain a text semantic vector corresponding to the search text.

[0581] In some embodiments, the text encoder of the CLIP model can be trained using text training samples. The training process of the text encoder of the CLIP model can be seen in the following embodiments.

[0582] S514: The multimodal understanding module 530 returns the text semantic vector corresponding to the search text to the search module 520.

[0583] S515: The search module 520 performs vector recall in the index library based on the text semantic vector corresponding to the search text.

[0584] Vector recall refers to recalling the index in the index library that matches the text semantic vector corresponding to the search text.

[0585] In some embodiments, vector similarities can be calculated between the text semantic vector corresponding to the search text and the visual semantic vectors included in multiple indexes in the index library to obtain vector similarity calculation results corresponding to the multiple indexes. N indexes with higher vector similarities among the multiple vector similarity calculation results are used as vector recall results. Where N is an integer greater than 0. For example, N can be a pre-set number of vector recall results, such as 5, 8, or 10.

[0586] Vector similarity refers to the degree of similarity between two vectors and can be calculated using a variety of methods. For example, the degree of similarity can be determined by calculating the cosine similarity of the two vectors. Other methods may also be used, and this application does not limit this.

[0587] In some embodiments, as described above, the vector space in the index library includes multiple vector regions, and each vector region includes multiple indexes with high vector similarity. In an embodiment of the present application, the distance between the text semantic vector corresponding to the search text and the cluster center points of multiple vector regions in the index library can be calculated to determine the cluster center point with the closest distance. Then, the vector similarity between the visual semantic vector and the text semantic vector respectively included in the multiple indexes in the vector region to which the cluster center point belongs is calculated. The indexes are sorted in order from large to small according to the vector similarity to obtain the inverted zipper corresponding to the cluster center point. Exemplarily, the first N indexes in the inverted zipper can be used as the vector recall result. Exemplarily again, the indexes whose vector similarity exceeds the vector similarity threshold can be used as the vector recall result.

[0588] The process of vector recall is described in detail below with reference to FIG7 .

[0589] As shown in Figure 7, the inverted index library includes multiple cluster centers, including cluster center 1. First, the distances between the text semantic vector corresponding to the search text and the cluster centers of multiple vector regions in the index library are calculated, and cluster center 1 is determined to be the closest cluster center. The vector similarities between the visual semantic vectors and the text semantic vectors of the multiple indexes in the vector region to which center 1 belongs are then calculated. These vectors are sorted from highest to lowest by vector similarity to obtain inverted zipper 1 corresponding to cluster center 1. The top N indexes from inverted zipper 1 corresponding to cluster center 1 are selected as the vector recall result.

[0590] In inverted zipper 1, the visual semantic vector corresponding to index 1 is closest to cluster center 1, while the distances between indexes 2 and 3 and cluster center 1 gradually increase. Therefore, selecting the TopN indexes is to select from index 1 and work backwards. N can be any integer greater than 0 and is not limited in this application.

[0591] S516: The search module 520 calls the natural language understanding module 540 to identify entities in the search text.

[0592] The search module 520 calls the natural language understanding module 540 to identify entities contained in the search text.

[0593] For example, named entity recognition (NER) can be used to identify entities with specific meanings in the search text. Entities can include, but are not limited to, time, place, person, organization, and proper nouns. For example, if the search text is "scenery taken last week," the entities in the search text include "last week" and "scenery."

[0594] S517: The natural language understanding module 540 returns the entities in the search text to the search module 520.

[0595] S518: The search module 520 performs entity recall in the index library based on the entities in the search text.

[0596] Entity recall refers to recalling the index in the index library that matches the entity in the search text.

[0597] In some embodiments, the index may include relevant information representing frames in the video segment and attribute labels of the video segment, which include entities. For example, the time when the video was acquired, the location where the video was acquired, the classification labels of the video segment, etc. may all include entities. Taking the search text "scenery taken in city B last week" as an example, there is an entity "city B" for the location, an entity "last week" for the time, and an entity "scenery" related to the content of the video. Then, a match can be made among the entities corresponding to the multiple indexes, and an index that matches the entity in the search text can be obtained as the entity recall result.

[0598] S519: The search module 520 sorts the vector recall results and the entity recall results.

[0599] In some embodiments, the intersection results or union results of the vector recall results and the entity recall results may be sorted.

[0600] For example, the sorting can be performed based on the vector similarity between the vector recall results and the search text, and the entity matching between the entity recall results and the search text. For example, the vector similarity between the text semantic vector of the search text and the visual semantic vector of the recall results (vector recall results or entity recall results), as well as the matching between the entities in the search text and the entities in the recall results, can be weighted and summed to obtain the comprehensive matching corresponding to the recall results. The recall results (including vector recall results and entity recall results) can then be sorted in descending order based on the comprehensive matching corresponding to each of the results.

[0601] In this way, based on the search text entered by the user, on the basis of the search matching of the text semantic vector, the search matching of the entities in the search text is performed, and the final displayed result order is obtained based on the comprehensive matching degree of the search results, ensuring that the video presented to the user is a result that is more matched with the user's search text, further improving the user experience.

[0602] S520 : The search module 520 returns the search results to the gallery service module 510 .

[0603] The search module 520 returns the sorted search results to the gallery service module 510 .

[0604] S521: The gallery service module 510 displays the search results to the user.

[0605] It is understood that in actual applications, the mobile phone 300 usually stores pictures and videos in the gallery application. Therefore, when searching in the search interface provided in the gallery application, both picture search results and video search results are displayed. In other words, the index library includes not only the index corresponding to the video segment, but also the index corresponding to the picture.

[0606] In some embodiments, the index corresponding to the picture may include but is not limited to the picture semantic vector, the attribute label of the picture, etc. For example, as described above, a video frame can be regarded as an image, and a picture is also an image, so the picture semantic vector corresponding to the picture can be generated by the image encoder of the CLIP model provided by the multimodal understanding module 530, and returned to the gallery service module 510. For example, the attribute label of the picture can be obtained by the gallery service module 510 first receiving and storing the user's addition or modification operation on the attribute label of the picture. The search module 520 then receives the picture semantic vector and the attribute label of the picture sent by the gallery service module 510, and constructs the index corresponding to the picture based on this.

[0607] For example, as shown in the first search result display interface 330 of FIG3 a , the first search result includes a video search result and an image search result.

[0608] In some embodiments, search results can be displayed based on multiple sorted indexes. As described above, the index corresponding to a video includes a visual semantic vector representing the frame, related information representing the frame, and attribute labels of the video segment. The index corresponding to an image includes an image semantic vector and attribute labels of the image.

[0609] For example, as shown in the first search result display interface 330 of Figure 3a, the first display area 331 of the image search results displays the image, while the video search results display thumbnails and time points. The information corresponding to the thumbnails and time points is the information included in the index corresponding to the video search results. Taking "Video A" in Figure 3a as an example, its thumbnail is the starting frame of video segment a included in the index, its first time point is the time point "02:18" corresponding to the starting frame of video segment a included in the index, and its second time point is the total duration of the video included in the index, "08:32".

[0610] As another example, the video search results and the image search results can also display their corresponding attribute tags. For example, the video search results can display the video shooting time and video shooting location. As shown in (2) in Figure 3b, the image search results include the shooting time of the image "October 1, 2023" and the shooting location of the image "City B".

[0611] As another example, the video search results may also display other content in the corresponding index, such as character relationships, names, etc. This application does not limit this.

[0612] It should be noted that the gallery service module, search module, multimodal understanding module, and natural language understanding module may also be located in the cloud server. That is, the cloud server utilizes the interaction of these four modules to implement the steps included in the index construction phase. During the search phase, the steps included in the search phase may be implemented based on the interaction between an electronic device such as mobile phone 300 and the cloud server.

[0613] In some embodiments, the mobile phone 300 can send the search text entered by the user to the gallery service module of the cloud server, so that the gallery service module of the cloud server interacts with other modules to implement the steps included in the search stage. The gallery service module of the cloud server then sends the search results to the mobile phone 300, so that the mobile phone 300 displays the search results to the user.

[0614] In some embodiments, assume that the total length of video E is 2 minutes, and it is shot at a frame rate of 30, that is, 30 images are shot in one second. When video E is stored in the cloud server, the multimodal understanding module in the cloud server can perform frame splitting processing on video E, decomposing video E into individual video frames, and 3600 video frames can be obtained; the multimodal understanding module then segments video E, segmenting it with 1s as the time unit, and dividing it into 120 video segments, each of which includes 30 video frames (that is, 30 images shot in one second); the multimodal understanding module scores the 30 video frames in each video segment, and uses the video frame with the highest score as the representative frame of the video segment. For example, scoring can be based on the jitter, clarity, and pixels of the video frame, which is not limited in this application.

[0615] The cloud server's search module then uses the visual semantic vector of the representative frame as the index of the corresponding video segment, allowing it to subsequently match the search text with the representative frame of each video segment of Video E. If the search text successfully matches the 50th video segment of Video E, the video thumbnail returned to the user will be the representative frame of the 50th video segment, and the return time is 50 seconds. The specific implementation of the above embodiment can be found in Figure 5b and the description of the above embodiment, and will not be repeated here.

[0616] It should be noted that the multimodal understanding module 530 of the mobile phone 300 can also use 1s as the time unit to segment the video stored in the gallery, and this application does not limit this.

[0617] For example, FIG8 is a schematic diagram showing the principle of another visual media search method provided in an embodiment of the present application.

[0618] The specific implementation process of the two stages is described below.

[0619] 1. Index construction phase

[0620] FIG9 is a flowchart of an example of a visual media search method provided by an embodiment of the present application. As shown in FIG9 , the index building phase of the method may include:

[0621] S101 . In response to a user adding and / or modifying visual media, the gallery APP stores the incremental visual media added and / or modified by the user and its attribute information in a gallery database.

[0622] Optionally, users can add new visual media by photographing, downloading, or taking screenshots. Furthermore, users can modify existing visual media to create new ones. These modifications include, but are not limited to, beautification, editing, customizing the addition of names, and adding watermarks.

[0623] The attribute information of the visual media may include, but is not limited to, identification information, acquisition location, acquisition time, name, and label (also known as classification label). Optionally, the identification information of the visual media may be represented by an identity document (ID) (e.g., a hash ID), or by a name, path, etc., which is not limited in this application. In the following embodiments, the identification information of the visual media is represented by an ID as an example.

[0624] The acquisition location refers to the location where the visual media was acquired. The acquisition time refers to the time when the visual media was acquired. For a captured video or image, the acquisition location may be the capture location, and the acquisition time may be the capture time. In one embodiment, for a captured video or image, the attribute information may also include information about the camera used for the capture (e.g., whether the capture was taken with a front-facing camera). For a screenshot, the acquisition location may be the screenshot location, and the acquisition time may be the screenshot time. For a downloaded video or image, the acquisition location may be the download location, and the acquisition time may be the download time.

[0625] A name is a user-added name for a person in an image or video. A tag is information that characterizes the type or attributes of an image or video. For example, a tag can be a person, plant, animal, building, or natural scenery.

[0626] For the target visual media and its attribute information added by the user, the Gallery APP will store the newly added visual media and its attribute information in the Gallery database. For the visual media and its attribute information modified by the user, the Gallery APP will synchronously modify the visual media and its attribute information stored in the Gallery database.

[0627] In other embodiments, the gallery APP can store the newly added visual media and its attribute information in the cloud to reduce the local storage pressure of the electronic device.

[0628] For ease of description, in the following embodiments, the visual media added and / or modified by the user is referred to as incremental visual media or target visual media, the video in the visual media added and / or modified by the user is referred to as incremental video or target video, and the image in the visual media added and / or modified by the user is referred to as incremental image or target image.

[0629] S102. The Gallery APP synchronizes the incremental visual media and its attributes to the CV database of the CV service.

[0630] When users add and / or modify visual media, the Gallery APP can synchronize the incremental visual media and its attributes to the CV database, which makes it easier for subsequent modules of the CV service to obtain data as needed.

[0631] S103: When the electronic device is in a charging state and the screen is in an off state, the gallery APP sends a segmentation request to the video segmentation module of the CV service.

[0632] Segment requests are used to request segmentation of incremental videos and obtain representative frames of the video segments. Optionally, a segment request can include the incremental video or its ID.

[0633] Optionally, the Gallery APP can subscribe to the charger plug-in status from the relevant module in the electronic device, and when the charger plug-in status changes, the module sends a notification to the Gallery APP. In this way, the Gallery APP can know in real time whether the charger is in the plugged-in state, that is, whether the electronic device is in the charging state. Similarly, the Gallery APP can also subscribe to the screen status from the relevant module in the electronic device, and when the screen status changes, the module sends a notification to the Gallery APP. In this way, the Gallery APP can know in real time whether the screen is in the off state.

[0634] In another embodiment, the gallery app can also send a segmentation request to the video segmentation module when the electronic device is charging, the screen is off, and the current time is within a preset time period. Optionally, the preset time period can be, for example, from 00:00 to 07:00 daily. This can reduce user interruptions and improve the user experience.

[0635] S104 : The video segmentation module segments the incremental video in response to the segmentation request to obtain multiple video segments and their attribute information, and extracts representative frames from each video segment.

[0636] Specifically, the video segmentation module can use a preset video segmentation algorithm to segment each incremental video. Video segmentation algorithm video segmentation is mainly divided into three steps: the first step is to deframe the incremental video, splitting the incremental video into multiple frame images; the second step is to segment the multiple frame images obtained by defragging based on the video semantics. Each segmented video segment can represent a scene, and the frames in each video segment have a certain correlation with each other; the third step is to extract a frame image from each video segment as a representative frame.

[0637] The representative frame is used to represent the semantics of a video segment. In one embodiment, the first frame, the last frame, or an intermediate frame of a video segment can be used as the representative frame. In another embodiment, the representative frame can also be determined by calculating parameters such as the jitter, clarity, label, and pixel transformation of each frame in the video segment. The present embodiment does not specifically limit the method for determining the representative frame.

[0638] Optionally, the attribute information of the video segment may include video segment identification information, start time, end time, label union, etc. The identification information of the video segment may be, for example, the video segment ID. The label union refers to the union of the labels (referred to as frame labels) of each frame image in the video segment.

[0639] The specific implementation of step S104 will be further described in subsequent embodiments.

[0640] S105 , the video segmentation module returns the video segments, their attribute information, and representative frames to the gallery APP.

[0641] S106. The gallery APP stores the video segment and its attribute information and representative frames in the gallery database.

[0642] S107. The gallery APP sends a face analysis request to the face analysis module in the CV service.

[0643] The face analysis request is used to request face analysis on incremental images containing faces (hereinafter referred to as incremental face images or target face images) and representative frames containing faces (hereinafter referred to as face representative frames). Face analysis can include face recognition, face clustering, etc.

[0644] Optionally, the face analysis request carries the incremental face image and the representative face frame. Alternatively, the face analysis request may carry identification information of the incremental face image and the representative face frame.

[0645] S108. The face analysis module performs face analysis on the incremental face image and the face representative frame in response to the face analysis request to obtain a face analysis result, which includes a face clustering result.

[0646] Specifically, the face analysis module can perform face recognition on the incremental face image and the face representative frame using a face recognition algorithm. Optionally, the face information can be represented by a face ID.

[0647] Afterwards, on the one hand, the face analysis module can analyze the gender, age, etc. of the person based on the recognized face data.

[0648] On the other hand, the face analysis module clusters the facial information in the recognized incremental face images and representative face frames together with the historical face information based on the face clustering algorithm to obtain multiple clusters, thereby updating the face clustering results obtained by the historical time period analysis.

[0649] Among them, historical faces refer to faces recognized in images within a historical time period. It is understandable that historical face information can be stored in the CV database. The face analysis module can perform face clustering based on the face recognition results of incremental face images and face representative frames, combined with the historical face information in the CV database. Face clustering is to cluster faces with similar features into a class, so that at least one class is obtained, and each class can correspond to a person. Optionally, the face analysis module can identify each person, for example, the characters can be numbered and an ID corresponding to each character can be set. In other words, each person can be represented by a character ID. The character ID is also called tag_id.

[0650] It should be noted that when the face analysis module performs face clustering, it may be unable to cluster, that is, for a certain face, there are no other faces with similar features. In this case, the face clustering algorithm can set the character ID corresponding to the face to a preset identifier, for example, to -1. In other words, a character ID of -1 indicates that there is no corresponding class. During the subsequent algorithm operation, characters with a character ID of the preset identifier can be eliminated and not included in the calculation scope. Only characters with a character ID other than the preset identifier are considered. This can improve the accuracy of character relationship recognition and enhance the user experience.

[0651] For example, Table 5 is an example of face clustering results provided in an embodiment of the present application.

[0652] Table 5

[0653] Taking the clustering results for numbers 9, 10, 11, and 12 in Table 5 as an example, we can see that the image with the image ID "26fe24ebd66f90b540cc79637efdc46d" contains four faces, whose corresponding face IDs are 0, 1, 2, and 3, respectively. Clustering these four faces yields the classes to which they belong, and the person IDs corresponding to these four face IDs are "ser_136254524879543," "ser_246762133257616," "ser_112952351578583," and "ser_275636865923451," respectively.

[0654] It should be noted that Table 5 is only an example of a face clustering result and does not constitute any limitation to the present application. In actual applications, the face clustering result may include more or less content than Table 5, or may be in other forms of expression.

[0655] S109. The face analysis module returns the face analysis results to the gallery APP.

[0656] S110 , the face analysis module stores the incremental face image and the face representative frame in the CV database, and updates the face image set in the CV database.

[0657] It is understood that each time the face analysis module receives a face analysis request, it saves the incremental face images and representative face frames requested for analysis, as well as their attribute information, into the CV database. This allows the CV database to store all images and representative face frames from the image library database. These images or representative frames can be stored in a collection. This facilitates subsequent processing such as character relationship analysis based on the face image collection. Images in a face image collection are also referred to as face images.

[0658] In addition, after each face analysis, the face analysis module also stores the face analysis results (including face clustering results, character age, character gender, etc.) in the CV database, which also facilitates subsequent character relationship analysis and other processing.

[0659] S111. The gallery APP stores the face analysis results in the gallery database.

[0660] S112. The gallery APP sends a character relationship analysis request to the character relationship analysis module in the CV service.

[0661] The person relationship analysis request is used to analyze the person relationship information of incremental facial images and representative face frames. This person relationship information characterizes the relationship type between the people in the image and the central person. Examples of person relationships include family, son, daughter, parent, colleague, and friend. The central person is the person at the center of the social relationship. This central person is generally considered to be the device owner or someone with whom the device owner has a close relationship.

[0662] Optionally, the person relationship analysis request may carry identification information of the incremental facial image and the representative facial frame.

[0663] S113. The character relationship analysis module responds to the character relationship analysis request and performs character relationship analysis based on the face clustering result and the face image set in the CV database to obtain character relationship information of the incremental face image and the face representative frame.

[0664] Character relationship analysis is also called social relationship analysis, character relationship discovery, etc.

[0665] Optionally, the character relationship analysis module can, on the one hand, obtain character relationship information that users actively add to characters. On the other hand, the character relationship analysis module can call a social circle recognition algorithm to identify the central person among the characters included in the face image collection based on the face clustering results, and identify the social circles formed by the characters in the face image collection based on the central person and the type of social circle. A character's social circle refers to a collection of characters with social relationships. The type of social circle is used to characterize the relationship between the characters in the social circle and the central person. Optionally, the type of social circle can include family, son, daughter, parents, colleagues, friends, etc. In other words, the type of social circle can serve as character relationship information.

[0666] S114. The character relationship analysis module returns the character relationship information of the incremental face image and the face representative frame to the gallery APP.

[0667] It is understood that each time the person relationship analysis module performs person relationship analysis based on the face clustering results and the face image collection, it obtains person relationship information for all images in the face image collection. The person relationship analysis module can save this person relationship information and return the person relationship information corresponding to the incremental face images and representative face frames to the Gallery app.

[0668] S115. The gallery APP stores the character relationship information of the incremental face image and the face representative frame in the gallery database.

[0669] S116. The gallery APP sends an image semantic understanding request to the multimodal understanding module in the CV service.

[0670] Image semantic understanding requests are used to request semantic understanding of incremental images and representative frames.

[0671] Optionally, the image semantic understanding request may carry an incremental image and a representative frame, or identification information of the incremental image and the representative frame.

[0672] S117. The multimodal understanding module performs semantic understanding on the incremental image and the representative frame in response to the image semantic understanding request to obtain a visual semantic vector.

[0673] Specifically, the multimodal understanding module can include a pre-trained image semantic understanding model. Each incremental image and representative frame is input into the image semantic understanding model to obtain the corresponding semantic vector. In this embodiment, the semantic vector corresponding to the image is called a visual semantic v...

Claims

1. A visual media search method, characterized in that: The method comprises: Displaying a first interface of the gallery application; the first interface includes an input box, and the input box includes a first text; The first interface includes a first thumbnail of the first video, and the first thumbnail includes a first time point; In response to a user triggering operation on the first thumbnail, playing the first video starting from the first time point; Displaying a second interface of the gallery application; the second interface includes the input box, and the input box includes a second text; The second interface includes a second thumbnail of the first video, and the second thumbnail includes a second time point; In response to the user's triggering operation on the second thumbnail, the first video is played starting from the second time point; the second time point is later than the first time point.

2. The method according to claim 1, characterized in that The first text is different from the second text, the first video matches the first text, and the first video matches the second text.

3. The method according to claim 2, characterized in that The first video includes a first video frame and a second video frame; The first video matches the first text, and the first video matches the second text, including: The first text matches a first video frame corresponding to the first thumbnail, and the second text matches a second video frame corresponding to the second thumbnail, where the second video frame is after the first video frame.

4. The method according to claim 3, characterized in that The step of playing the first video starting from the first time point in response to a user triggering an operation on the first thumbnail includes: Play the first video starting from the first video frame or the first starting frame; The step of playing the first video from the second time point in response to the user triggering the second thumbnail includes: Play the first video starting from the second video frame or the second starting frame; The first start frame is before the first video frame, the second start frame is before the second video frame, and the second start frame is after the first video frame.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Displaying a third interface of the gallery application; the third interface includes the input box, the input box includes third text, the third text is different from the first text, and the third interface includes the first thumbnail; In response to the user's triggering operation on the first thumbnail, the first video is played starting from the first time point.

6. The method according to any one of claims 1 to 4, characterized in that The first interface further includes a third thumbnail of the first video, and the third thumbnail includes a third time point; the method further includes: In response to the user's triggering operation on the third thumbnail, the first video is played starting from the third time point; the first text matches the third video frame corresponding to the third thumbnail; the third time point is later than the second time point.

7. The method according to claim 6, characterized in that The playing of the first video from the third time point includes: The first video is played starting from the third video frame or the third start frame; the third start frame is before the third video frame, and the third start frame is after the second video frame.

8. The method according to any one of claims 1 to 7, characterized in that The first interface further includes a fourth thumbnail of the second video, the fourth thumbnail including a fourth time point, and the method further includes: In response to the user's triggering operation on the fourth thumbnail, the second video is played starting from the fourth time point; the first text matches the fourth video frame corresponding to the fourth thumbnail, and the second video is different from the first video.

9. The method according to claim 1, characterized in that Also includes: receiving a first operation of a user inputting a fourth text through the search box; In response to the first operation, a plurality of hit results matching the fourth text are searched and obtained, wherein the plurality of hit results include at least one hit image and at least one hit video segment, and the at least one hit video segment is a video segment in at least one target video; Determining a matching score for each of the hit results, wherein the matching score represents a degree of matching between the hit result and the fourth text; Sorting the at least one hit image and the at least one target video according to the matching scores of the hit results to obtain a first sorting result; According to the first sorting result, the at least one hit image and the at least one target video are displayed.

10. The method according to claim 9, characterized in that The step of sorting the at least one hit image and the at least one target video according to the matching scores of the hit results to obtain a first sorting result includes: For a first target video, the highest score among the matching scores of all first hit video segments is used as the matching score of the first target video; the first target video is any one of the at least one target video, and the first hit video segment is the hit video segment in the first target video; The at least one hit image and the at least one target video are sorted according to the matching scores to obtain the first sorting result.

11. The method according to claim 9, characterized in that The number of the first hit video segments is multiple, and the method further includes: The plurality of first hit video segments are sorted according to the matching scores to obtain a second sorting result.

12. The method according to claim 11, characterized in that After displaying the at least one hit image and the at least one target video according to the first sorting result, the method further includes: receiving a second operation of the user to play the first target video; In response to the second operation, the multiple first hit video segments are played in sequence according to the second sorting result.

13. The method according to any one of claims 9 to 11, characterized in that The displaying of the at least one hit image and the at least one target video includes: Displaying thumbnails of the hit images and thumbnails of the target videos; The thumbnail of the first target video is a thumbnail of a frame of image in the second hit video segment, and the second hit video segment is a video segment with the highest matching score among all the first hit video segments; The thumbnail of the first target video displays one or more of the following data: the start playing time of the second hit video segment, the end playing time of the second hit video segment, the middle playing time of the second hit video segment, the time, the total number of the first hit video segments, and the total number of video segments contained in the first target video.

14. A visual media search method, characterized in that: The method comprises: Displaying a first user interface of a gallery application; the first user interface includes thumbnails of a plurality of image resources in the gallery application, the image resources including pictures or videos; In response to a sliding gesture of the user on the first user interface, the thumbnails of the plurality of image resources move in a direction of the sliding gesture, and a first button is displayed on the first user interface; In response to a user clicking operation on the first button, a search interface is displayed; the search interface includes a first search box, the first search box is used to receive text information input by the user, and the text information is used to search for image resources.

15. The method according to claim 14, characterized in that The method further comprises: After the user stops the sliding gesture on the first user interface, a second user interface is displayed; the second user interface includes a second search box.

16. The method according to claim 15, characterized in that The method further comprises: If a next sliding gesture is not detected within a first period of time after a sliding gesture is detected, it is determined that the sliding gesture performed by the user on the first user interface has stopped.

17. The method according to claim 14 or 15, characterized in that The method further comprises: In response to a user clicking operation on the second search box, the search interface is displayed.

18. The method according to any one of claims 14 to 17, characterized in that: The first user interface of the display gallery application includes: displaying a third user interface of the gallery application, wherein the third user interface includes thumbnails of a plurality of image resources in the gallery application; In response to a sliding gesture of the user on the third user interface, a first user interface of the gallery application is displayed; the image resources included in the first user interface are different from the image resources included in the third user interface.

19. The method according to claim 18, characterized in that The third user interface includes the second search box, and the method further includes: In response to a sliding gesture of the user on the third user interface, hiding the second search box.

20. The method according to any one of claims 14 to 19, characterized in that: The method further comprises: In response to receiving a first text message input by a user in the first search box, displaying a first search result interface; the first search result interface includes a first thumbnail of the first video, the first thumbnail including a first time point; In response to receiving the second text information entered by the user in the first search box, a second search result interface is displayed; the second search result interface includes a second thumbnail of the first video, the second thumbnail includes a second time point, and the second time point is later than the first time point.

21. The method according to claim 20, characterized in that The image of the first thumbnail is an image frame corresponding to the first time point in the first video, and the image of the second thumbnail is an image frame corresponding to the second time point in the first video.

22. The method according to claim 20 or 21, characterized in that The method further comprises: In response to a user clicking on the first thumbnail, playing the first video starting from the first time point; In response to a user clicking operation on the second thumbnail, the first video is played starting from the second time point.

23. The method according to claim 20 or 21, characterized in that The method further comprises: In response to a user clicking on the first thumbnail, a first playback interface of the first video is displayed; the first playback interface includes multiple marking points, and the marking points are used to indicate the starting position of the video segment corresponding to the first text information in the first video.

24. The method according to any one of claims 20 to 23, characterized in that: The first text information is different from the second text information. The first video includes a first image frame and a second image frame. The first text information matches the first image frame corresponding to the first thumbnail. The second text information matches the second image frame corresponding to the second thumbnail. The second image frame is after the first image frame.

25. An electronic device, characterized in that: Includes memory, display, and processor; The memory is coupled to the processor, the memory is used to store computer program code, the computer program code includes computer instructions, the display screen provides a display function, and the one or more processors call the computer instructions to enable the electronic device to execute the visual media search method as described in any one of claims 1-13, or the visual media search method as described in any one of claims 14-24.

26. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the visual media search method according to any one of claims 1 to 13, or the visual media search method according to any one of claims 14 to 24.