Video searching method, electronic equipment and storage medium
By using CLIP model to search videos in electronic devices, the problem of inability to effectively handle complex search texts in the prior art is solved, and the accuracy of video search and user experience are improved.
Patent Information
- Application Number
- CN202311544834.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-30
AI Technical Summary
When searching for videos, the existing electronic devices' gallery cannot effectively process the complex search text entered by the user, resulting in the inability to accurately search for videos that match the user's needs. Users need to manually search, which is complicated to reduce the user experience.
By implementing a video search method in an electronic device, the CLIP model is used to match the search text input by the user with the visual semantics of the video frame, display the video thumbnail corresponding to the user's needs, and allow the user to start playing video from different points in time to meet the user's various search needs.
It improves the accuracy of video search results, enables electronic devices to accurately search and display videos that match user needs, reduces user operation steps and improves user experience.
Smart Images

Figure CN120066355A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a video search method, an electronic device, and a storage medium. Background Art
[0002] With the rapid development of computer technology, the use of electronic devices such as mobile phones, tablet computers, or laptop computers has entered a popularization stage. Users are accustomed to using electronic devices to shoot videos and store them in the gallery application (hereinafter simply referred to as the gallery) of the electronic device, or download videos of interest on the Internet and store them in the gallery of the electronic device, so that they can view, edit, or share the stored videos with others at any time in the gallery of the electronic device.
[0003] The gallery of an electronic device may contain multiple videos. To facilitate users' search, the gallery of the electronic device usually supports a search function. For example, a video may have attribute tags such as shooting time or shooting location. When a user enters search text such as time or location in the search box of the gallery, the electronic device can display videos that match the search text entered by the user.
[0004] However, the gallery only supports simple search text matching such as attribute tags. When a user enters complex search text, the electronic device may not be able to search for videos that exactly correspond to the user's needs, and the user still needs to manually search through them, which is cumbersome and greatly reduces the user experience. Summary of the Invention
[0005] A video search method, an electronic device, and a storage medium provided by this application aim to improve the accuracy of video search results, enable the electronic device to accurately search for and display videos corresponding to the user's needs, reduce user operations, and enhance the user experience.
[0006] To achieve the above object, this application adopts the following technical solutions:
[0007] In a first aspect, the present application provides a video search method, which is applied to an electronic device. The electronic device can be a device including a gallery application such as a mobile phone, a tablet computer, a laptop computer, etc. The video search method includes: The electronic device displays a first interface of the gallery application. For example, the first interface can be an album display interface, a photo display interface, a time point display interface, a creation display interface, etc. of the gallery application. The first interface includes an input box, which supports the function of searching for pictures / videos in the gallery application. The input box in the first interface includes a first text, and the first text is the search text input by the user, such as "a girl doing yoga", "a boy standing near the bridge", etc. The first interface includes a first thumbnail of a first video, and the first thumbnail corresponds to a video frame in the first video. The first thumbnail includes a first time point, and the first time point is a time stamp of the first video. When the user performs a trigger operation on the first thumbnail, such as a click operation, etc. The electronic device starts playing the first video from the first time point.
[0008] The electronic device displays a second interface of the gallery application. The second interface includes an input box, and the input box includes a second text. The second interface includes a second thumbnail of the first video, and the second thumbnail includes a second time point, and the second time point is later than the first time point. When the user performs a trigger operation on the second thumbnail. The electronic device starts playing the first video from the second time point. That is, after the user triggers the first thumbnail or the second thumbnail, the electronic device plays the first video, but the starting time points of the playback are different, that is, the video frames shown to the user are different.
[0009] Based on the first text input by the user, the electronic device can search for the first video. After the user performs a trigger operation on the first thumbnail, the electronic device can play the first video from the first time point, and the picture content played from the first time point meets the user's needs; Based on the second text input by the user, the electronic device can search for the first video. When the user triggers the second thumbnail, the electronic device starts playing the first video from the second time point later than the first time point, and the played picture content meets the user's needs. That is, after the user inputs the text, the video search results displayed by the electronic device can meet the user's needs. In this way, the electronic device can accurately display the video corresponding to the user's needs, reduce the user's operations, and thus improve its user experience.
[0010] In a possible implementation, the first text entered by the user in the search box is different from the second text. For example, the first text can be "a boy standing by the river", and the second text can be "a boy standing near the bridge". The first video matches the first text. For example, the first video includes picture content about "a boy standing by the river". The first video also matches the second text. For example, the first video includes picture content about "a boy standing near the bridge". In this way, based on the search text entered by the user, the electronic device can display videos that match the search text, meeting the user's needs without the user having to manually search, thus enhancing the user experience.
[0011] In a possible implementation, the first video includes a first video frame and a second video frame. The second video frame comes after the first video frame. That is, when playing the first video, it will first play to the first video frame and then to the second video frame. The first thumbnail corresponds to the first video frame. That is, the displayed first thumbnail is the first video frame of the first video. The second thumbnail corresponds to the second video frame. That is, the displayed second thumbnail is the second video frame of the first video. The first text matching the first video further includes: the first text matching the first video frame. For example, the first text is "a boy standing by the river", and the picture content of the first video frame will have a boy standing by the river. The second text matching the first video further includes: the second text matching the second video frame. For example, the second text is "a boy standing near the bridge", and the picture content of the second video frame will have a boy standing near the bridge.
[0012] In this way, the search text entered by the user matches the picture content shown in the video frame corresponding to the thumbnail. The search text reflects the user's needs. Therefore, the electronic device can search for videos that exactly match the user's needs, improving the accuracy of video search and thus enhancing the user experience.
[0013] In a possible implementation, when the user triggers the first thumbnail of the first video and the electronic device starts playing the first video from the first time point, the electronic device can start playing the first video from the first video frame or the first starting frame, and the first starting frame is before the first video frame. That is, the electronic device can start playing from the first video frame that matches the first text, or start playing from the first starting frame and then play to the first video frame that matches the first text.
[0014] Similarly, when the user triggers the second thumbnail of the first video and the electronic device starts playing the first video from the second time point, the electronic device can start playing from the second video frame that matches the second text, or start playing from the second starting frame before the second video frame, and can play to the second video frame, and the second starting frame is after the first video frame.
[0015] After the user triggers the thumbnail, the electronic device can start playing from a video frame that matches the search text, or can start playing from a previous video frame. At the same time, the second starting frame is before the second video frame and after the first video frame, and the first video frame and the second video frame match different first text and second text respectively. This means that if the user triggers the second thumbnail, the first video frame that matches the first text will not be shown to the user, but instead, the electronic device can start playing from the second starting frame that is closer to the second video frame. Compared with the first video frame, the content shown in the second starting frame has a stronger correlation with the content shown in the second video frame. In this way, the content shown in the video played to the user can either have stronger coherence or more accurately correspond to the user's needs, further enhancing the user's experience.
[0016] In a possible implementation manner, the video search method further includes: The electronic device displays a third interface of the gallery application, and the third interface includes an input box, and the input box includes third text. The third text is different from the first text. For example, the first text can be "a boy standing by the river", and the third text can be "a boy standing outdoors". The third interface includes a first thumbnail of the first video. When the user performs a trigger operation on the first thumbnail, the electronic device starts playing the first video from the first time point.
[0017] In practical applications, different users may have different text descriptions for the same video, and the same user may also change the text description for the same video. That is, the user enters different search texts, but may want to search for the same video search result. In this application, the first text can search for the first video, and the third text can search for the first video. The electronic device displays the first thumbnail in both cases. That is, the electronic device can search for the same video search result based on different search texts. In this way, the electronic device can search for the same video that accurately corresponds to the user's needs based on different search texts, which can improve the accuracy of video search and enhance the user's experience.
[0018] In a possible implementation manner, the first interface further includes a third thumbnail of the first video, and the third thumbnail includes a third time point. The video search method further includes: When the user triggers the third thumbnail, the electronic device starts playing the first video from the third time point. The first text matches the third video frame corresponding to the third thumbnail, and the third time point is later than the second time point. It was mentioned above that the first text can also match the first video frame corresponding to the first thumbnail. That is, the same search text can search for different video frames of the same video. In this way, different video frames of the video can be searched based on the same search text. Therefore, this application can display search results that accurately correspond to the search text based on the search text, improving the accuracy of video search and further enhancing the user's experience.
[0019] In a possible implementation, when the user triggers the third thumbnail of the first video and the electronic device starts playing the first video from the third time point, the electronic device can start playing from the third video frame that matches the first text, or start playing from the third starting frame before the third video frame, so that the electronic device can play up to the third video frame, and the third starting frame is after the second video frame.
[0020] In this way, the first video can be started from the third starting frame that has a stronger correlation with the third video frame, or directly from the third video frame. In both cases, the picture content that matches the first text can be presented to the user, meeting the user's needs and enhancing the user experience.
[0021] In a possible implementation, the first interface further includes a fourth thumbnail of the second video. The fourth thumbnail includes a fourth time point. The video search method further includes: when the user triggers the fourth thumbnail, the electronic device starts playing the second video from the fourth time point. The first text matches the fourth video frame corresponding to the fourth thumbnail. In this way, based on the same search text input by the user, different video frames of different videos that match the search text can be searched, improving the accuracy of video search and further enhancing the user experience.
[0022] In a possible implementation, the video search method further includes: the electronic device displays a negative first screen interface. The negative first screen interface includes a search box that supports online search and search of local files of the electronic device. The search box includes the first text. The negative first screen interface includes a first thumbnail of the first video. The first thumbnail includes a first time point. When the user triggers the first thumbnail, the electronic device can start playing the first video from the first time point.
[0023] In this way, based on the support for the user to search for videos through the search box of the gallery application, the present application also supports the user to search for videos through the search box of the negative first screen interface, further enhancing the user experience.
[0024] In a second aspect, the present application provides a video search method applied to an electronic device. The electronic device can be a device including a gallery application such as a mobile phone, a tablet computer, a laptop computer, etc. The video search method includes: The electronic device displays a first interface of the gallery application. For example, the first interface can be an album display interface, a photo display interface, a time point display interface, a creation display interface, etc. of the gallery application. The first interface includes an input box, which supports the function of searching for pictures / videos in the gallery application. The input box in the first interface includes a first text, which is the search text input by the user, such as "a boy standing outdoors", "a girl dancing indoors", etc. The first interface includes a first thumbnail of a first video. The first thumbnail includes a first time point, which is a timestamp in the first video. The first text is matched with the first video frame corresponding to the first thumbnail through the CLIP model, that is, the text semantics of the first text is matched with the visual semantics of the first video frame. The CLIP model is a pre-trained neural network model for matching images and texts, that is, it can match the first text input by the user and the first video frame. When the user performs a trigger operation on the first thumbnail, such as a click operation. The electronic device starts playing the first video from the first time point.
[0025] The electronic device displays a second interface of the gallery application, which includes an input box, and the input box includes a second text. The second interface includes a second thumbnail of the first video, and the second thumbnail includes a second time point, which is later than the first time point. The second text is matched with the second video frame corresponding to the second thumbnail through the CLIP model, that is, the text semantics of the second text is matched with the visual semantics of the second video frame. When the user performs a trigger operation on the second thumbnail. The electronic device starts playing the first video from the second time point, that is, after the user triggers the first thumbnail or the second thumbnail, the electronic device plays the first video, but the starting time points of the playback are different, that is, the video frames shown to the user are different.
[0026] In this way, based on the search text input by the user, video frames matching the search text can be searched. The text semantics of the search text is fully associated with the visual semantics of the video frame, enabling the fusion interaction between the search text and the picture content shown in the video frame, thereby improving the accuracy of video search and enhancing the user experience.
[0027] In a possible implementation, dividing a video can obtain some video segments. The first video frame is in the first video segment of the first video, and the second video frame is in the second video segment of the first video. The first video frame and the second video frame can be determined through the following steps: The electronic device first performs a first process on the first video. For example, the first video is frame-split to obtain multiple video frames of the first video and the classification label of each video frame. The classification label refers to the type of the object shown in the video frame. For example, the classification label can be a person, a plant, an animal, a building, or a natural scenery, etc. The electronic device then performs a second process on the first video based on the classification labels respectively corresponding to the multiple video frames. For example, the first video is segmented, and the first video is divided into multiple video segments, obtaining multiple video segments including the first video segment and the second video segment. Subsequently, the electronic device determines a video frame of the first video segment as the first video frame, that is, determines the first video frame as the representative frame of the first video segment, which can be used to represent the first video segment; the electronic device determines a video frame of the second video segment as the second video frame, and determines the second video frame as the representative frame of the second video segment, which can be used to represent the second video segment.
[0028] In this way, if this solution is implemented on an electronic device, there is no need to build an index for each video frame of the video, avoiding increasing the latency required for search and matching. Instead, an index of the first video segment can be built based on the first video frame, and an index of the second video segment can be built based on the second video frame, greatly reducing the number of indexes, improving the speed of video search, further enhancing the user experience, and also saving the cost consumed when building the index.
[0029] In a possible implementation, when the electronic device performs the second process on the first video based on the classification labels respectively corresponding to the multiple video frames, it further includes: The electronic device performs the second process on the first video based on the difference in the image parameters respectively corresponding to the adjacent video frames of the multiple video frames of the first video, and based on the difference in the classification labels respectively corresponding to the adjacent video frames of the multiple video frames. The image parameter is the display characteristic of the video frame, such as the jitter degree, the clarity, the pixel value, etc. For example, the electronic device can perform the second process on the first video based on the difference in the clarity respectively corresponding to the adjacent video frames such as the first video frame and the second video frame of the first video, and the difference between the classification labels respectively corresponding to the adjacent video frames such as the first video frame and the second video frame. Thus, multiple video segments including the first video segment and the second video segment are obtained.
[0030] In this way, based on the classification labels, the types of objects shown in multiple video frames can be determined, and based on the image parameters, the changes in the picture quality shown in multiple video frames can be determined. Video frames with more similar displayed picture content can be grouped into a video segment, which is beneficial to determining the video frames that can represent the video segment, so that the representative frames corresponding to multiple video segments of a video can more completely represent the picture content shown in the video.
[0031] In a possible implementation, the electronic device determines a video frame of the first video segment as the first video frame, that is, the representative frame that can represent the first video segment. Further, it includes: The electronic device can determine the first starting frame of the first video segment, that is, the first video frame, the first ending frame, that is, the last video frame, a random video frame among the multiple video frames of the first video segment, or the video frame corresponding to the time point at the first position as the first video frame. The time point at the first position can be obtained based on the average value of the starting time point and the ending time point of the first video segment, that is, the middle time point of the first video segment, and the video frame corresponding to the middle time point is the middle frame of the first video segment. In this way, any video frame of the video segment can represent the video segment, which is convenient for the electronic device to build an index for the video segment.
[0032] In a possible implementation, the electronic device determines a video frame of the first video segment as the first video frame. Further, it includes: The electronic device determines the first video frame based on the image parameters corresponding to the multiple video frames of the first video segment and the classification labels corresponding to the multiple video frames of the first video segment. Based on the picture quality of the video frame and the types of objects shown in the video frame, a video frame that is more representative for the first video segment is determined. In this way, the accuracy of the index established based on the representative frame can be improved, the accuracy of video search can be further improved, and the user experience can be enhanced.
[0033] In a possible implementation, the matching of the first text and the first video frame corresponding to the first thumbnail through the CLIP model further includes: The electronic device inputs the first text into the text encoder of the CLIP model to obtain the text semantic vector of the first text, which can represent the semantic features of the entire first text. The electronic device then inputs the first video frame into the image encoder of the CLIP model to obtain the visual semantic vector of the first video frame, which can represent the semantic features of the first video frame. Subsequently, the electronic device matches the first text and the first video frame based on the text semantic vector and the visual semantic vector.
[0034] In this way, the first video frame can be used to represent the first video segment, so its visual semantic vector is fully associated with the first video segment. Based on the text semantic vector of the first text, the electronic device matches the visual semantic vector of the first video frame, which can achieve the fusion interaction between the search text and the content shown in the representative frame, thereby improving the accuracy of video search and enhancing the user experience.
[0035] In a possible implementation, the first vector similarity between the text semantic vector of the first text and the visual semantic vector of the first video frame is greater than or equal to the first threshold. In this way, by setting the first threshold, the visual semantic vector of the first video frame with a high vector similarity is determined, and then a more accurate video search result is displayed to the user.
[0036] In a possible implementation, the inverted index library includes multiple visual semantic vectors. The electronic device can perform clustering on them to determine multiple clustering centers and the clustering clusters respectively corresponding to the multiple clustering centers. The visual semantic vector of the first video frame belongs to the first clustering cluster corresponding to the first clustering center. Therefore, the second vector similarity between the text semantic vector of the first text and the vector of the first clustering center in the inverted index library is greater than or equal to the second threshold. The third vector similarity between the text semantic vector of the first text and the visual semantic vector is greater than or equal to the third threshold.
[0037] In this way, when the electronic device performs search matching based on the text semantic vector of the search text, it can first match it with multiple clustering centers, and then match it with the visual semantic vectors of the clustering clusters of the determined clustering centers, avoiding matching with all the indexes in the index library. Further avoiding search latency, improving the video search efficiency, and thus enhancing the user experience.
[0038] In a possible implementation, the entity of the first text matches the entity of the attribute label of the first video segment. In this way, based on the search text input by the user, the electronic device not only performs search matching of the text semantic vector but also performs search matching of the entity in the search text. That is, the electronic device can perform vector recall and entity recall, and can display to the user a video search result that is both a vector recall result and an entity recall result, further enhancing the accuracy of video search and the user experience.
[0039] In a possible implementation, the first interface displayed by the electronic device further includes a third thumbnail of a second video, and the third video frame corresponding to the third thumbnail is in the third video segment of the second video. The entity of the first text matches the entity of the attribute label of the third video segment. In this way, on the basis of vector recall, the electronic device can also perform entity recall, and can display the entity recall result and the vector recall result to the user, improving the richness of the video search result, further improving the accuracy of the video search, and enhancing the user experience.
[0040] In a possible implementation, the first interface displayed by the electronic device further includes a fourth thumbnail of a first video, the first text matches the fourth video frame corresponding to the fourth thumbnail, the fourth video frame is in the fourth video segment of the first video, and the entity of the first text matches the entity of the attribute label of the fourth video segment. Moreover, the entity of the first text matches the entity of the attribute label of the first video segment, and the first text matches the first video frame. On the first interface displayed by the electronic device, the display order of the first thumbnail is before the fourth thumbnail.
[0041] The display order of the first thumbnail and the fourth thumbnail can be determined through the following steps: The electronic device determines the first comprehensive matching degree between the first thumbnail and the first text based on the first vector similarity between the visual semantic vector of the first video frame and the text semantic vector of the first text, and the first matching degree between the attribute label of the first video segment and the entity of the first text. The electronic device then determines the second comprehensive matching degree between the second thumbnail and the first text based on the second vector similarity between the visual semantic vector of the fourth video frame and the text semantic vector of the first text, and the second matching degree between the attribute label of the fourth video segment and the entity of the first text. The electronic device then displays the first thumbnail before the fourth thumbnail according to the order from largest to smallest of the comprehensive matching degrees, that is, the first comprehensive matching degree is greater than the second comprehensive matching degree, and the thumbnail with a higher comprehensive matching degree is placed in the front for display.
[0042] In this way, sorting the video search results based on the comprehensive matching degree ensures that the video search results in better positions are results that are more matched with the user's search text, further enhancing the user experience.
[0043] In a third aspect, the present application provides an electronic device, which includes a memory, a display screen, and a processor; the memory stores computer program code, and the computer program code includes computer instructions; the display screen provides a display function; one or more processors call the computer instructions to enable the electronic device to execute the method of the first aspect or the second aspect above.
[0044] Fourthly, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method according to the first aspect or the second aspect is implemented.
[0045] As can be seen from the above technical solutions, the present application has the following beneficial effects:
[0046] Based on the first text input by the user, the first video can be searched and the first thumbnail is displayed. After the user triggers it, the first video can be played starting from the first time point included in the first thumbnail, and the picture content played by the electronic device starting from the first time point can meet the user's needs. Based on the second text input by the user, the first video can be searched and the second thumbnail is displayed. After the user triggers it, the first video can be played starting from the second time point included in the second thumbnail, and the picture content played by the electronic device starting from the second time point can meet the user's needs. That is, the first video is matched based on both the first text and the second text, but under the user's trigger operation, the start time point of playing the first video is different. Thus, based on the text input by the user, the matching picture content can be displayed to the user to meet the user's needs, which can reduce the user's manual searching operation, thereby improving the user's usage experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1a It is a schematic diagram of a search result display interface provided by an embodiment of the present application;
[0048] Figure 1b It is a schematic diagram of another search result display interface provided by an embodiment of the present application;
[0049] Figure 2a It is a schematic diagram of the composition example of an electronic device provided by an embodiment of the present application;
[0050] Figure 2b It is a schematic diagram of the software structure example of an electronic device provided by an embodiment of the present application;
[0051] Figure 3a It is an application scenario of a video search method provided by an embodiment of the present application;
[0052] Figure 3b It is a schematic diagram of a video playback interface provided by an embodiment of the present application;
[0053] Figure 3c It is a schematic diagram of a search interface provided by an embodiment of the present application;
[0054] Figure 4a It is an application scenario of another video search method provided by an embodiment of the present application;
[0055] Figure 4b Another application scenario of the video search method provided by the embodiments of the present application;
[0056] Figure 5a A schematic diagram of a video search method provided by the embodiments of the present application;
[0057] Figure 5b A signaling interaction diagram of a video search method provided by the embodiments of the present application;
[0058] Figure 6 A schematic diagram of the determination process of a representative frame provided by the embodiments of the present application;
[0059] Figure 7 A schematic diagram of a vector recall process provided by the embodiments of the present application;
[0060] Figure 8 A schematic diagram of a CLIP model training process provided by the embodiments of the present application;
[0061] Fig. 9 A schematic diagram of another CLIP model training process provided by the embodiments of the present application. Detailed implementation manners
[0062] Next, in combination with related technologies, the technical advantages of a video search method provided by the present application will be compared and described. For the convenience of understanding, an example scenario will be used for illustration. In this example scenario, the electronic device is a mobile phone, and multiple pictures and multiple videos are stored in the gallery application of the mobile phone.
[0063] First, the vocabulary involved in the embodiments of the present application will be described. It can be understood that this description is for a clearer understanding of the embodiments of the present application and does not necessarily constitute a limitation on the embodiments of the present application.
[0064] Video frame: It refers to any frame of a video. A frame is a still picture in a video, and consecutive frames can form a video.
[0065] Video segment: It refers to the segment obtained by dividing a video. In some embodiments, the video can be first frame-split to obtain individual video frames of the video, and then the video segment algorithm can be used to perform segmentation processing on the video including multiple video frames to obtain video segments. The specific implementation manner can be referred to the description of the following embodiments.
[0066] Representative frame: It is a video frame in a video segment, which can be used to represent the video segment. Exemplarily, the representative frame can be the starting frame, the ending frame, a randomly selected video frame, or the optimal frame with the highest score, etc.
[0067] CLIP Model: The CLIP (Contrastive Language-Image Pre-Training) model is a pre-trained neural network model for matching images and text. In some embodiments, through contrastive learning based on the text encoder (TextEncoder) and image encoder (Image Encoder) of the CLIP model, a text encoder for outputting text semantic vectors of text can be trained, as well as an image encoder for outputting visual semantic vectors of images or video frames.
[0068] Text Semantic Vector: It can be obtained by inputting text into the text encoder, and it is a vector that can represent the semantic features of the entire text. Exemplarily, the text encoder can adopt models such as Transformer commonly used in Natural Language Processing (NLP), and this application does not make limitations in this regard. In the embodiments of this application, the search text input by the user can be input into the text encoder to obtain the text semantic vector of the search text.
[0069] Visual Semantic Vector: It can be obtained by inputting an image or a video frame of a video into the image encoder. Exemplarily, the image encoder can adopt a CNN model or a VIT model, and this application does not make limitations in this regard. In the embodiments of this application, the representative frame of the video can be input into the image encoder to obtain the visual semantic vector of the representative frame.
[0070] Vector Similarity: It is used to describe the similarity between two vectors (for example, between a text semantic vector and a visual semantic vector). In the embodiments of this application, by comparing the similarity between the text semantic vector of the search text and the visual semantic vector, the video frame that matches the search text can be determined. Exemplarily, the vector similarity can be calculated through the cosine similarity calculation formula. Of course, it can also be calculated through other methods.
[0071] Entity: A word with a specific meaning in the text. Exemplarily, entities can include but are not limited to time, place, person name, organization name, and proper noun in the text. In some embodiments, named entity recognition technology (NER) can be used to identify entities with specific meanings in the text, and this application does not make limitations in this regard.
[0072] Image Parameter: It is used to indicate the display characteristics of an image or a video frame. Exemplarily, image parameters can include the jitter degree, clarity, pixel value, etc. of the video frame, and this application does not make limitations in this regard.
[0073] Attribute label: Information used to indicate the attributes of a video or a video segment. In the embodiments of the present application, the attribute labels of a video may include the video acquisition location, the video acquisition time, a person's name, a file name, or a classification label, etc.
[0074] In the related art, when the gallery of an electronic device stores a video, it also stores the shooting location or shooting time of the video and keywords in the file name of the video as attribute labels, and uses the attribute labels as the index of the video. Subsequently, the user can input simple search texts such as time, location, or keywords in the file name in the search interface of the gallery, and the search matching between the search text input by the user and the index (attribute label) of the video can be performed to achieve video search.
[0075] Suppose the electronic device is a mobile phone 300, and a video 1 is stored in its gallery, and the video 1 has been pre-assigned attribute labels of "stars" and "this year". Exemplarily, "stars" can be the file names manually configured by the user for the video 1. "This year" is the video acquisition time of the video 1. As Figure 1a shown, when the search text input by the user in the search box of the gallery provided by the mobile phone 300 is the keyword "stars", the mobile phone 300 can search for the video with the attribute label of "stars" and display the search results including the video 1.
[0076] However, as Figure 1b shown, if the user inputs the complex search text of "a boy standing beside a tree" in the search box of the gallery, the mobile phone 300 cannot search for the video 1 and still requires the user to manually rummage through.
[0077] In practical applications, the user may input a complex search text such as "a boy standing beside a tree" based on the content shown in the video to describe the video they want to obtain. However, in the related art, the electronic device only supports search matching of keywords such as attribute labels, which easily causes the electronic device to be unable to display the video required by the user, and still requires the user to manually operate the scroll bar of the gallery to rummage through, with cumbersome operations, thereby affecting the user experience of using the mobile phone 300.
[0078] To solve the above problems, the embodiments of the present application provide a video search method, which can be applied to an electronic device. For the convenience of understanding, the composition and software structure of the electronic device are introduced.
[0079] This application does not limit the type of electronic devices. For example, the electronic device can be a mobile phone, a tablet computer, a desktop type, a laptop, a notebook computer, an Ultra-mobile Personal Computer (UMPC), a handheld computer, a netbook, a Personal Digital Assistant (PDA), a wearable electronic device, a smart watch, etc. This application does not impose special restrictions on the specific forms of the above-mentioned electronic devices.
[0080] In this embodiment, as Figure 2a shown, the electronic device may include a processor 110, an internal memory 120, a camera 130, a display screen 140, an audio module 150, a speaker 150A, and a headphone jack 150B.
[0081] The processor 110 may include one or more processing units. For example, the processor 110 may include a video codec and / or a neural-network processing unit (NPU), etc.
[0082] The processor 110 may also be provided with a memory for storing instructions and data.
[0083] The internal memory 120 may be used to store computer-executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory 120.
[0084] The internal memory 120 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function (such as the sound playback function during video playback, the image playback function during video playback, etc.). The data storage area may store data created during the use of the electronic device (such as video data, etc.).
[0085] In some embodiments, the instructions stored in the internal memory 120 are for executing a video search method. The processor 110 may implement video search by executing the instructions stored in the internal memory 120.
[0086] In some embodiments, the electronic device performs a video search in the videos stored in the gallery application. The videos stored in the gallery application can be those shot and stored by the user using the electronic device. The electronic device can implement the shooting function through an ISP, the camera 130, a video codec, a GPU, the display screen 140, and an application processor, etc.
[0087] The ISP is used to process the data fed back by the camera 130. In some embodiments, the ISP may be disposed in the camera 130. The camera 130 is used to capture still images or videos. The video codec is used to compress or decompress digital videos. In this way, the electronic device can play or record videos in a variety of coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0088] The NPU can be used to implement applications such as the intelligent cognition of the electronic device, such as: image recognition, face recognition, speech recognition, text understanding, etc. In some embodiments, the NPU can be used to perform text understanding on the search text input by the user in the search interface provided by the electronic device.
[0089] The electronic device realizes the display function through the GPU, the display screen 140, and the application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 140 and the application processor.
[0090] A series of graphical user interfaces (GUIs) can be displayed on the display screen 140 of the electronic device, and these GUIs are all the main screens of the electronic device. Generally speaking, the display screen 140 of the electronic device includes limited controls, and the user can interact with the controls through direct manipulation, so as to read or edit the relevant information of the application program.
[0091] In some embodiments, the electronic device may include a gallery application. The display screen 140 of the electronic device can display the icon corresponding to the gallery application. After being triggered by the user, the display screen 140 of the electronic device can display the search interface of the gallery application, and the search interface includes a search control. The user can perform an editing operation on the search control to input the search text, and perform a triggering operation, so that the electronic device can search in the gallery application based on the search text input by the user, and display the search results to the user through the display screen 140.
[0092] In some embodiments, if the user triggers the video stored in the gallery application, the electronic device can realize the audio function during video playback through the audio module 150, the speaker 150A, the headphone jack 150B, and the application processor, etc.
[0093] The audio module 150 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The electronic device can listen to the audio during video playback through the speaker 150A. The headphone jack 150B is used to connect a wired headphone. The electronic device can listen to the audio during video playback through the wired headphone connected to the headphone jack 150B.
[0094] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may include more or fewer components than shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0095] In addition, on top of the above components, the electronic device also runs an operating system. For example, the iOS operating system developed by Apple Inc., the Android open-source operating system developed by Google Inc., the Windows operating system developed by Microsoft Corporation, etc. Application programs can be installed and run on this operating system.
[0096] The operating system of the electronic device can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In this application embodiment, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the electronic device.
[0097] Figure 2b It is the software structure block diagram of the electronic device in this application embodiment.
[0098] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0099] The application layer may include a series of application packages. As Figure 2b shown, the application packages may include application programs such as the gallery service module, the camera, music, and video playback applications.
[0100] In some embodiments, the gallery service module stores pictures / videos obtained under operations such as user shooting, downloading, screenshotting, or screen recording; is also used to store representative frames of video segments, visual semantic vectors of representative frames, and other information; also receives search text input by the user so that the electronic device can perform video search and matching based on the search text input by the user; and also displays pictures / videos obtained under operations such as user shooting, downloading, screenshotting, or screen recording.
[0101] In some embodiments, the video playback application may be the native video playback application of the electronic device. In some embodiments, the video playback application may also be a third-party video playback application.
[0102] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. For example, Figure 2b as shown, the application framework layer may include a window manager, a content provider, a resource manager, a view system, etc.
[0103] The window manager is used to manage window programs. The content provider is used to store and obtain data and enable these data to be accessible to applications. The data may include videos, images, etc. The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, and so on. The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications.
[0104] Android Runtime includes a core library and a virtual machine. Android runtime is responsible for the scheduling and management of the Android system. The core library contains two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core library of Android.
[0105] The system library may include functional modules such as a surface manager, Media Libraries, a 3D graphics processing library (e.g., OpenGL ES), a 2D graphics engine (e.g., SGL), etc.
[0106] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications. The Media Libraries support the playback and recording of multiple common audio and video formats, as well as static image files, etc. The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc. The 2D graphics engine is the drawing engine for 2D drawing.
[0107] The kernel layer is the layer between hardware and software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.
[0108] In addition, the electronic device may also include functional modules such as a search module, a multimodal understanding module, and a natural language understanding module. These modules may all be located in the same layer of the electronic device, may be located in different layers of the electronic device respectively, or may be located in multiple layers of the electronic device at the same time to implement their functions through the software interfaces between layers.
[0109] In some embodiments, the search module is used to build an index corresponding to a video based on the information stored in the above-mentioned gallery service module, and is also used to perform vector recall among multiple indexes based on the text semantic vector corresponding to the search text, and perform entity recall among multiple indexes based on the entities in the search text. It is also used to sort the vector recall results and entity recall results to obtain search results.
[0110] In some embodiments, the multimodal understanding module is used to perform frame splitting on a video to obtain individual video frames, and then use a video segmentation algorithm to perform segmentation processing on the video to obtain video segments. Subsequently, representative frames of the video segments and visual semantic vectors of the representative frames are determined.
[0111] In some embodiments, the natural language understanding module is used to identify entities in the search text input by the user.
[0112] The video search method provided by the embodiments of the present application can be implemented based on the interaction among the four modules: the gallery service module, the search module, the multimodal understanding module, and the natural language understanding module. For the specific implementation method, reference can be made to Figure 5b and the detailed introduction in the following embodiments.
[0113] It should be noted that although the embodiments of the present application are described by taking the Android system as an example, its basic principle also applies to electronic devices based on operating systems such as iOS and Windows.
[0114] To enable those skilled in the art to more clearly understand the technical solution of the present application, the application scenario of the technical solution of the present application will be described first below.
[0115] Exemplarily, the video search method provided by the embodiments of the present application can be implemented in the gallery search scenario of an electronic device.
[0116] Combined with Figure 3a As shown, taking the electronic device as the mobile phone 300 as an example, the video search method provided by the embodiments of the present application will be exemplarily described below. In this scenario, the gallery of the mobile phone 300 may include multiple pictures and multiple videos. In some embodiments, the pictures / videos may be taken by the user through the mobile phone 300 and stored in the gallery, or may be downloaded and stored in the gallery by the user through other application platforms, or may be downloaded and stored in the gallery after the user sends other electronic devices to the mobile phone 300.
[0117] In this video search method, the mobile phone 300 can build an index for the videos stored in the picture gallery when it is charging or the screen is off. During the process of building the index, the mobile phone 300 can perform frame splitting on the video to obtain individual frames, use a video segmentation algorithm to segment the video based on multiple video frames to obtain video segments, then determine the representative frames of the video segments and the visual semantic vectors of the representative frames, and subsequently build an index for the video segments based on the visual semantic vectors of the representative frames. For the specific implementation method, reference can be made to Figure 5b and the detailed introduction in the following embodiments. It should be understood that the pictures stored in the picture gallery do not need to be frame-split, and the index of the pictures can be built by referring to the method of building an index based on the representative frames of video segments.
[0118] Suppose the user took multiple videos during a weekend outing using the mobile phone 300 and stored them in the mobile phone 300. When the user has free time and hopes to edit the videos taken during the outing, the user can search through the search function of the picture gallery of the mobile phone 300 at this time, so that the mobile phone 300 displays the videos taken during the outing.
[0119] As Figure 3a shown, the mobile phone 300 displays the main interface 310, and the main interface 310 includes application icons 311 of multiple application programs such as the application icon 311 of the picture gallery. When the user triggers the application icon 311 of the picture gallery, the mobile phone 300 starts the picture gallery in response to the user's trigger operation and displays the album display interface 320 of the picture gallery. The album display interface 320 includes a search box 321 and multiple albums. For example, the multiple albums can include Figure 3a the "All Photos" album, "Camera" album, "My Favorites" album, "Screenshots & Screen Recordings" album, "My Favorites" album, "Self-created" album, and "Video Editing" album shown, and so on.
[0120] In one example, the user can perform a trigger operation on the search box 321 included in the album display interface 320 and enter the search text "boys standing outdoors" in the search box 321. The mobile phone 300 converts the "boys standing outdoors" entered by the user into a text semantic vector; then, based on the text semantic vector, it performs a search match with the indexes of videos and pictures. The index of the video includes the visual semantic vectors of the representative frames, and the index of the picture includes the visual semantic vectors of the pictures. Then, the vector similarity between the text semantic vector and the visual semantic vector can be calculated, and the pictures / videos corresponding to the visual semantic vectors whose vector similarity exceeds the vector similarity threshold are used as the first search results. Subsequently, the mobile phone 300 can display the first search result display interface 330 of the picture gallery, and the first search result display interface 330 can display the first search results corresponding to "boys standing outdoors", and the first search results can include pictures and videos.
[0121] The index of the video in the above first search result includes the visual semantic vector of the representative frame. As introduced above, the representative frame can be used to represent the corresponding video segment. Therefore, the visual semantics of the representative frame can represent the visual semantics of multiple video frames of the video segment, that is, the visual semantic vector of the representative frame is fully associated with the video segment. When the user inputs a complex search text based on the picture content shown in the video, the present application performs a search match between the text semantic vector corresponding to the search text and the visual semantic vector of the representative frame, which can achieve the fusion interaction between the search text and the picture content shown in the video segment, improve the accuracy of video search, and thus enhance the user experience.
[0122] In some embodiments, the first search results displayed on the mobile phone 300 can be sorted in descending order according to the degree of relevance to the search text input by the user.
[0123] As Figure 3a shown, the first search result display interface 330 of the mobile phone 300 may include a first display area 331, a second display area 332, and a third display area 333. In the first display area 331, all the first search results, including pictures and videos, are displayed. In the second display area 333, a most matching video search result is displayed, and the number "4" of all video search results is shown, and the time point corresponding to the most matching video search result is enlarged and displayed. In the third display area 332, a most matching picture search result is displayed, and the number "6" of all picture search results is shown.
[0124] As Figure 3a shown, the video search result of the first search result display interface 330 shows a thumbnail and a time point. In some embodiments, the thumbnail corresponds to a video frame of a video segment in the video, and the visual semantic vector of the representative frame in the video segment matches the text semantic vector corresponding to the search text input by the user, that is, the vector similarity between the two exceeds the vector similarity threshold. The time point of the video can be the time point of the video segment to which the video frame corresponding to the thumbnail belongs. Exemplarily, the time point can be the time point of the video frame corresponding to the thumbnail, or the time point of other video frames in the video segment.
[0125] Taking the video segment a of a video search result "Video A" in the first display area 331 as an example, "Video A" is divided into multiple video segments. The thumbnail shown in "Video A" corresponds to the starting frame of the video segment a, and the vector similarity between the visual semantic vector of the representative frame of the video segment a and the text semantic vector of "a boy standing outdoors" exceeds the vector similarity threshold.
[0126] The second display area 332 includes two time points. The first time point is "02:18", indicating that the starting frame of video segment a starts to be displayed at "02:18", and the second time point indicates that the total duration of "Video A" is "08:32". As Figure 3a shown, if "Video A" is the video required by the user, the user can trigger "Video A", and the mobile phone 300 can respond to the user's trigger operation on "Video A" and display the video playback interface 340. In this interface, the mobile phone 300 adjusts the time progress of "Video A" to "02:18" and starts playing, that is, starts playing from the starting frame of video segment a in "Video A".
[0127] It should be noted that the thumbnail and time points shown in the above "Video A" are only examples.
[0128] For example, the thumbnail can correspond to the representative frame of video segment a in "Video A". The representative frame can be the starting frame, ending frame, middle frame, optimal frame or any one of the video frames of video segment a, etc. It is also possible to use the starting frame of video segment a as the thumbnail, and the present application does not make any limitations in this regard.
[0129] For example, the first time point can be the same as the time point of the video frame corresponding to the thumbnail. For example, if the thumbnail corresponds to the representative frame of video segment a, the first time point can be the time point corresponding to the representative frame; the first time point can also be different from the time point corresponding to the thumbnail. For example, the thumbnail is the representative frame of video segment a, the representative frame is the middle frame of video segment a, and the first time point can be the time point corresponding to the starting frame of this video segment a. The present application does not make any limitations in this regard.
[0130] It should be noted that the above method of displaying the search box 321 through the album display interface 320 for the user to input search text is only an example, and the search box can also be displayed through other interfaces of the gallery.
[0131] In some embodiments, as Figure 3a shown, at the bottom of the album display interface 320, there are "Photo" control, "Time Point" control and "Creation" control, which are used to enter the photo display interface, time point display interface and creation display interface of the gallery respectively. The photo display interface, time point display interface and creation display interface all include a search box. The search boxes of these three interfaces support the same search function as the search box 321 of the album display interface 320, and all support searching in the pictures and videos included in the gallery based on the search text input by the user. After the search is completed, the first search result display interface 330 is also displayed, and the included first search results are also the same.
[0132] It should be understood that if the vector similarity between the visual semantic vector of the representative frame of a video segment and the text semantic vector of a certain text exceeds a certain threshold, the video segment can be searched when the user inputs the text.
[0133] In some embodiments, a video includes multiple video segments, and the search text input by the user may match the index of one or more video segments of a video.
[0134] For example, the search text "a boy standing outdoors" input by the user can match video segments a, b, and c in "Video A". Then, video segments a, b, and c of "Video A" can be displayed simultaneously. The thumbnail shown for video segment b of "Video A" can be the representative frame of video segment b, and the thumbnail shown for video segment c of "Video A" can be the starting frame of video segment c. As Figure 3a shown in the first display area 331 of the first search result display interface 330, for the search text "a boy standing outdoors" input by the user, video segments a, b, and c of Video A can be searched. Video segments a, b, and c are sorted in the order of their matching degrees with the search text. This indicates that the vector similarities between the visual semantic vectors of the respective representative frames of these three video segments and the text semantic vector of "a boy standing outdoors" all exceed the vector similarity threshold.
[0135] In other embodiments, different search texts input by the user can all search to the same video segment. This will be introduced in combination with Figure 3b below.
[0136] As Figure 3b shown in (1) of, the first search result display interface 350 can include a first display area 351 and a second display area 352. Among them, all video search results are shown in the first display area 351. All picture search results are shown in the second display area 352, and the shooting time and shooting location of each picture search result can be shown on the right side of it.
[0137] Suppose the user enters "a boy standing by the river" in the input box. The first display area 351 can display the video search results that match it, including the video segment b of "Video B". The first display area 351 includes a thumbnail of the representative frame of the video segment b and two time points. The first time point is the time "05:48" corresponding to the starting frame of the video segment b, and the second time point represents the total duration of "Video B", which is "14:18". If the user triggers the thumbnail of the representative frame of the video segment b, the mobile phone 300 can display the video playback interface 360. In this interface, the mobile phone 300 adjusts the time progress of "Video B" to "05:48" and starts playing, that is, the mobile phone 300 starts displaying from the starting frame of the video segment b in "Video B".
[0138] Optionally, the first time point can also be the time point of the representative frame of the video segment b of Video B. This solution does not make any restrictions here.
[0139] It should be noted that Figure 3b In (1), it is only an example. The first display area 351 can also display other information of "Video B", such as the shooting location, shooting time, names of people, and relationships between people corresponding to "Video B". The second display area 352 can also display other information of the image search results, such as the names of people and relationships between people corresponding to the image.
[0140] It should be noted that after the user triggers the thumbnail of the representative frame of the video segment b, the mobile phone 300 adjusts the time progress of "Video B" to the first displayed time point and starts playing. It should be understood that when the first time point is the time point of the representative frame of the video segment b of Video B, if the middle frame of the video segment b is the representative frame, the time progress can also be adjusted to the time point corresponding to the middle frame and start playing. The time point corresponding to the middle frame is "07:09", which is later than the time "05:48" corresponding to the starting frame.
[0141] As Figure 3b As shown in (2), the first search result display interface 370 can include a first display area 371 and a second display area 372. Among them, all the image search results are displayed in the first display area 371. All the video search results are displayed in the second display area 372.
[0142] Suppose the user enters "a boy standing near the bridge" in the input box. The second display area 372 can display the video search results that match it, including the video segment c of "Video B" and the video segment b of "Video B". Combining Figure 3bAs shown in (1), for the two different search texts "a boy standing by the river" and "a boy standing near the bridge" input by the user, the video segment b of video "B" is retrieved. This indicates that the visual semantic vector of the representative frame of video segment b of "Video B" has a vector similarity exceeding the vector similarity threshold with both the text semantic vectors of "a boy standing by the river" and "a boy standing near the bridge". For the display form of video segment b of "Video B", refer to Figure 3b the introduction in (1), which will not be elaborated here.
[0143] The first display area 372 includes a thumbnail of the representative frame of this video segment c and two time points. The first time point is the time "08:32" corresponding to the representative frame of video segment c, and the second time point represents the total duration "14:18" of "Video B". If the user triggers the thumbnail of the representative frame of video segment c, the mobile phone 300 can display the video playback interface 380, in which the mobile phone 300 adjusts the time progress of "Video B" to "07:26" and starts playing, that is, the mobile phone 300 starts playing from the starting frame of video segment b in "Video B". Figure 3b For the display forms of the video search results and image search results in (2), refer to Figure 3a and Figure 3b the introduction in (1), which will not be elaborated here.
[0144] It should be understood that recommended information can be displayed in the search box of the gallery application, which is beneficial for the user to obtain relevant information about the pictures / videos stored in the gallery.
[0145] In some embodiments, the recommended information displayed in the search box of the gallery is a sentence with natural semantics. The user can input text in the search box according to the content and format of the recommended information, and use a sentence with natural semantics as the search text input into the search box. The electronic device searches for pictures / videos in the gallery according to the search text with natural semantics, and can more accurately search for the pictures / videos required by the user.
[0146] In one implementation, the electronic device can obtain the attribute tags of pictures / videos. In one example, the attribute tags include at least one of time, location, classification tag, and event.
[0147] In some embodiments, "time" can be obtained according to the shooting time of the picture / video. Exemplarily, "time" can include "Thursday, November 9, 2023", "Friday, July 2, 2021", etc.
[0148] "Location" can be obtained based on the shooting location of the picture / video. For example, it can be obtained according to the GPS positioning information when shooting the video. Exemplarily, "location" can include cities, scenic spots, etc.
[0149] "Classification label" and "event" can be obtained by performing semantic analysis on the picture / video. In one example, the electronic device can use computer vision services to perform semantic analysis on the video frames in the picture or video, and generate the content of "classification label" and "event" according to its semantics.
[0150] Exemplarily, "classification label" can include: people, plants or animals, etc., and can also include: buildings, natural scenery, etc., and can also include: street art, musical instruments, art exhibitions, sports competitions, birthdays, etc. For example, people can include portrait names, occupation names, person age group names, etc.
[0151] Exemplarily, "event" can include: games, sports, tourism, etc.
[0152] It should be understood that for a picture / video, its corresponding attribute labels can include one or more of time, location, classification label, and event. For example, for a picture downloaded from the network, if the electronic device fails to obtain its shooting time, then the attribute label does not include "time" or the content of "time" is empty.
[0153] In one implementation, by splicing one or more attribute labels of the picture / video according to a preset rule, at least one piece of combined information can be generated. Exemplarily, the electronic device splices the four attributes of the video, "time", "location", "classification label", and "event", and can generate a piece of combined information. Again exemplarily, the electronic device can also generate a piece of combined information according to one attribute label of the picture / video. For example, the attribute label is "time".
[0154] In one implementation, by splicing a fixed splicing word with a piece of combined information, the recommended content corresponding to the picture / video can be generated. Exemplarily, the fixed splicing word can include "try searching", "at", "shooting", "video", etc.
[0155] Table 1 shows some examples of the recommended content generated according to different numbers of attribute labels. It should be noted that when generating the recommended content, the content of the attribute labels can be mapped correspondingly. For example, map the specific time "Monday, October 1, 2023" to "today", "the day before yesterday", "August", "last month", "last year", "this year", or "National Day", etc.
[0156] Table 1
[0157]
[0158] As shown in Table 1, multiple spliced contents can be generated according to the splicing combination method shown in Table 1 for the attribute tags of a video. For example, if the attribute tags of a video include "category tag", "time", and "location", 5 different spliced contents can be generated respectively according to the splicing combination methods of serial numbers 1-5 in Table 1.
[0159] In one implementation, the electronic device can determine any one of the multiple spliced contents as the recommended content corresponding to the picture / video.
[0160] In another implementation, each of the above splicing combination methods corresponds to a priority. According to the order from front to back in Table 1, the priorities of the splicing combination methods decrease in turn. The electronic device generates the recommended content corresponding to the picture / video according to the splicing combination method with the highest priority among the splicing combination methods supported by the picture / video. As shown in Table 2, if the attribute tags of a video include "category tag", "time", and "location", then according to the splicing combination method of serial number 1 in Table 1, the fixed splicing words are spliced with "category tag", "time", and "location" to generate the recommended content corresponding to the picture / video.
[0161] In the video search method provided by the embodiments of the present application, the recommended information in the search box can be generated according to the recommended content corresponding to a picture / video in the picture library.
[0162] In one implementation, the electronic device can select the recommended content corresponding to a first picture / video in the picture library to generate the recommended information in the search box. In one example, it can be a first picture / video within a preset time period; for example, the preset time period is more than one month away from today. In one example, the electronic device updates the first picture / video once a day, and the first picture / videos selected within a preset duration (such as within a week) are not repeated.
[0163] Exemplarily, on the first day, the recommended information displayed in the search box of the electronic device's picture library is "Try to search for the video of Zhang San taking a walk by the river the day before yesterday"; on the second day, the recommended information displayed in the search box of the electronic device's picture library is "Try to search for the pictures of the seaside in August last year";... Within a preset duration (such as within a week), the first picture / videos selected by the electronic device every day are not repeated. Correspondingly, the recommended information displayed in the search box of the electronic device's picture library is not repeated every day.
[0164] Based on the above Figure 3a scene example, still taking the electronic device as the mobile phone 300 as an example, assuming that the user took multiple videos during the weekend outing and stored them in the picture library of the mobile phone 300. Then the mobile phone can generate the recommended information in the search box according to the recommended content corresponding to the video.
[0165] As shown Figure 3c in FIG. 390, the mobile phone 300 displays an album display interface 390, and the album display interface 390 includes a search box 321. In one example, the recommended information "Try to search for the video of Zhang San by the river last year" is displayed in the search box 321. This recommended information is used as an example of the search text input by the user. The user can input the search text in the search box according to the content and format of the recommended information for searching.
[0166] Combined with Figure 3c as shown in FIG. 301, after the user clicks to trigger the search box 321, the mobile phone 300 displays a search interface 301. The search interface 301 includes a search box 321 where the user can input text. When the mobile phone 300 detects the operation of the user inputting text in the search box 321, it can search in the picture library according to the search text in the search box 321.
[0167] For example, the recommended information "Try to search for the video of Zhang San by the river last year" is displayed in the search box 321. "Zhang San by the river last year" is a sentence with natural semantics. Compared with single tags such as "last year", "by the river", and "Zhang San", it is easier to search for the videos that the user really needs to find. For example, there are 7230 pictures / videos corresponding to the tag "last year" in the picture library, 1985 pictures / videos corresponding to the tag "by the river", and 1985 pictures / videos corresponding to the tag "Zhang San". The number of pictures / videos searched using single tags is relatively large. However, when searching in the picture library according to "Zhang San by the river last year", the number of pictures / videos searched will be significantly reduced. Thus, displaying sentences with natural semantics as recommended information is beneficial to providing a convenient and fast search experience for users.
[0168] Exemplarily, the video search method provided in the embodiments of the present application can be implemented in the global search scenario of an electronic device.
[0169] Combined with Figure 4a as shown in FIG. 300, still taking the electronic device as the mobile phone 300 as an example, the video search method provided in the embodiments of the present application will be exemplarily described below. In this scenario, the picture library of the mobile phone 300 may include multiple pictures and multiple videos. The sources of the pictures / videos can refer to the above examples and will not be elaborated here.
[0170] Combined with Figure 4a as shown in FIG. 300, based on the above Figure 3a scenario example, still assuming that the user took multiple videos during weekend outings and stored them in the picture library of the mobile phone 300 through the mobile phone 300. As shown Figure 4aAs shown, the mobile phone 300 displays the main interface 410. The user can perform a swiping operation to the right on the main interface 410, and the mobile phone 300 can display the minus one screen interface 420. The minus one screen interface 420 provides a global search function, which can provide users with rich online search services and the search service for the local resources of the mobile phone (i.e., the local files stored in the mobile phone 300).
[0171] The minus one screen interface 420 includes a search box 421. The user can enter the search text "boys standing outdoors" in the search box 421. The mobile phone 300 can perform an online search based on the entered search text and search for local resources in the local files included in the mobile phone 300. After the search is completed, the mobile phone 300 displays the second search result display interface 430, and the second search result display interface 430 can display the second search results corresponding to the entered "boys standing outdoors".
[0172] As Figure 4a shown, the second search result display interface 430 may include a first display area 431, a second display area 432, and a third display area 433. In some embodiments, the first display area 431 includes the second search results of the online search, the second display area 432 includes the second search results of the search in the application, and the third display area 433 includes the second search results of the local files.
[0173] In some embodiments, the second search results of the search in the application include the second search results of the search in the gallery application. Exemplarily, the content displayed in the second display area 432 is Figure 3a the content displayed in a partial area of the first search result display interface 330 in
[0174] It should be noted that the content displayed in the above-mentioned second display area 432 is only an example, and it can also display Figure 3a the content displayed in other areas of the first search result display interface 330 in
[0175] It should be noted that the above method of displaying the search box 421 through the minus one screen interface 420 for the user to enter the search text for global search is only an example. In some embodiments, the user can also perform a downward pull operation on the main interface of the mobile phone 300. The mobile phone 300 displays the main menu interface, and the main menu interface includes a search box, which can also provide the global search function, and the displayed search results are the same as the second search results.
[0176] Exemplarily, the video search method provided by the embodiments of the present application can be implemented in the search scenario of a video playback application on an electronic device.
[0177] In this scenario, the video search method is implemented through the interaction between the electronic device and the cloud server.
[0178] Still taking the electronic device as the mobile phone 300 as an example, in this scenario, the mobile phone 300 includes a video playback application.
[0179] As Figure 4b shown, the mobile phone 300 displays the main interface 440, and the main interface 440 includes application icons of multiple application programs such as the application icon 441 of the video playback application. The user triggers the application icon 441 of the video playback application, and the mobile phone 300 displays the first interface 450 of the video playback application. The first interface 450 includes a search box 451 and a search control 452. The user can enter the search text "documentary about city B" in the search box 451 and trigger the search control 452. In response to the user's triggering operation on the search control 452, the mobile phone 300 can search among multiple videos on the cloud server based on the search text input by the user. After the search is completed, the mobile phone 300 displays the third search result display interface 460 of the video playback application. The search result display interface includes multiple third search results, which are videos. The thumbnails and time points shown by them can be referred to the above example and will not be elaborated here.
[0180] In some embodiments, the third search result display interface 460 may further display the release time of the third search result, that is, the time when the third search result is stored on the cloud server. As Figure 4b shown, taking "Video C" included in the third search result display interface 460 as an example, the third search result display interface 460 also displays the release time "2018-05-03" of "Video C", indicating that "Video C" was stored on the cloud server on May 3, 2018. Exemplarily, it may be that the user uploaded "Video C" to the video playback application on May 3, 2018, so that the cloud server stores this "Video C".
[0181] In some embodiments, as Figure 4b shown, the third search result display interface 460 includes "Video C", and the first time point of "Video C" is "05:48". If "Video C" is the video required by the user, the user can trigger "Video C", and the mobile phone 300 can interact with the cloud server to start playing from "05:48" of "Video B".
[0182] Next, the video search method provided by the present application will be introduced in detail in combination with Figure 5a As Figure 5aAs shown, the video search method includes an index construction phase and a search phase.
[0183] In the index construction phase, in some embodiments, an index can be constructed for each video frame of the video.
[0184] In some embodiments, if the solution is implemented on a terminal device and an index is constructed for each video frame of the video, frame-by-frame matching is required in the search phase, which will increase the latency required for searching and reduce the user experience. Optionally, in some embodiments, the video can be segmented, and a video index can be constructed in units of segments.
[0185] Exemplarily, as Figure 5a shown, first, the video is frame-split to obtain multiple video frames of the video; then, a video segmentation algorithm is used to segment the video based on the multiple video frames of the video to obtain multiple video segments such as video segment 1 and video segment 2; subsequently, representative frames can be selected from the multiple video frames in the video segment, and a score evaluation is performed on the multiple video frames in the video segment respectively, and the video frame with the highest score is determined as the representative frame of each video segment; the representative frame is input into the image encoder of the CLIP model, and the visual semantic vector of the representative frame is output; an index is constructed based on the visual semantic vector corresponding to the representative frame, and an inverted index library is obtained based on the constructed index.
[0186] It should be noted that the above method for selecting representative frames is only an example, and the start frame, middle frame, end frame or a random video frame of the video segment can also be selected as the representative frame. The specific implementation can be seen in the introduction of the following embodiments. In the search phase, the user inputs a search text in the search interface; the search text is input into the text encoder of the CLIP model, and the text semantic vector corresponding to the search text is output; vector recall is performed from the inverted index library based on the text semantic vector corresponding to the search text; the search text is sent to the natural language understanding module; the natural language understanding module performs entity recognition on the search text to obtain the entities in the search text; entity recall is performed from the inverted index library based on the entities in the search text; based on vector recall and entity recall, the vector recall result and the entity recall result are returned from the inverted index library; the vector recall result and the search result are sorted to obtain the search result; the search result is displayed on the search interface.
[0187] In some embodiments, the image encoder of the CLIP model and the text encoder of the CLIP model can be trained by a training method based on contrastive learning.
[0188] Next Figure 5b Still taking the electronic device as a mobile phone 300 as an example, in combination with Figure 2b the gallery service module, search module, multimodal understanding module, and natural language understanding module shown in the system library, the video search method provided by the embodiments of the present application will be described in detail.
[0189] As Figure 5b shown, the video search method provided by the embodiments of the present application can be divided into the following two stages: an index construction stage and a search stage.
[0190] First, the steps included in the index construction stage will be introduced in detail in combination with Figure 5b the following.
[0191] S501: The gallery service module 510 receives an addition operation or a modification operation of the user for a video.
[0192] The addition operation of the video refers to the operation of the user storing the video in the gallery application. Exemplarily, the addition operation of the user for the video can be the operation of the user shooting a video through the mobile phone 300, downloading a video, or recording the screen of the mobile phone 300. The modification operation of the video refers to the operation of the user modifying the video already stored in the gallery application. Exemplarily, the modification operation of the user for the video can be the operation of the user cropping, splicing, adding special effects, or adding subtitles to the video, etc.
[0193] In some embodiments, the gallery service module 510 receiving an addition operation or a modification operation of the user for a video includes: the gallery service module 510 receiving an addition operation or a modification operation of the user for the attribute tags of the video.
[0194] Exemplarily, the attribute tags of the video can include the video acquisition location (such as the location where the video is shot, the download source of the video, etc.), the video acquisition time (such as the shooting time, the download time, or the screen recording time), the person name or classification tag, event, etc. newly added or modified by the user for the video. Among them, the classification tag can be used to indicate the type of the object shown in the video. Exemplarily, the classification tag can be a person, a plant, an animal, a building, or a natural scenery, etc. The event can be used to indicate what the object shown in the video does. Exemplarily, the event can be a game, a sport, etc.
[0195] In some embodiments, the classification tag can be manually configured by the user or automatically classified by the electronic device for the video.
[0196] S502: The gallery service module 510 stores the video.
[0197] In response to the addition operation or the modification operation of the user for the video, the gallery service module 510 stores the video in the mobile phone 300. Exemplarily, the video can be stored in the local file of the mobile phone 300, and the user can view the video through various channels such as the local folder of the mobile phone 300 and the gallery application. Also exemplarily, under the authorization operation of the user, the video can be stored in the cloud for backup to relieve the memory pressure of the mobile phone 300.
[0198] In some embodiments, in response to a user's operation of adding or modifying the attribute tags of a video, the gallery service module 510 may store the attribute tags of the video.
[0199] S503: The gallery service module 510 invokes the multimodal understanding module 530 to determine the representative frame of the video.
[0200] It can be understood that a video usually includes multiple video frames. Assuming that corresponding visual semantic vectors are determined for each video frame and stored as indexes, that is, a video corresponds to a large number of indexes, which will waste a large amount of computing resources and a large amount of storage space. At the same time, there may be noise information in a large number of indexes, and it also takes time to perform search matching, which is likely to affect the video search results.
[0201] Therefore, in the embodiments of the present application, the video is first segmented by the multimodal understanding module 530, and the corresponding representative frames are determined from each video segment, so as to determine the corresponding visual semantic vectors for the representative frames and store them as indexes in the subsequent steps, greatly reducing the number of indexes corresponding to the video. In this way, the representative frames corresponding to multiple video segments can more completely represent the video semantics of the video, and at the same time, costs can be saved.
[0202] In some embodiments, the gallery service module 510 may invoke the computer vision service provided by the multimodal understanding module 530 to determine the representative frame of the video. The computer vision service means that the multimodal understanding module 530 performs video semantic understanding on the video, and then the multimodal understanding module 530 determines the representative frame of the video.
[0203] Video semantic understanding means enabling the mobile phone to understand the meaning expressed by the content shown in the video, such as understanding information such as the types, quantities, positions of objects in the video, and the relationships between objects.
[0204] In some embodiments, the computer vision service provides services such as frame splitting processing of the video, segmenting the video using a video segmentation algorithm, and determining the representative frame of the video segment. Then the process of determining the representative frame of the video can be mainly divided into the following steps 1-step 3:
[0205] Step 1: The multimodal understanding module 530 uses the computer vision service to perform frame splitting processing on the video.
[0206] It can be understood that a video is composed of multiple video frames, and each video frame is a still picture in the video, that is, each video frame can be regarded as an image. Frame splitting processing means that the multimodal understanding module 530 uses the computer vision service to decompose the video into individual video frames.
[0207] In some embodiments, during the process of the multi-modal understanding module 530 performing frame splitting on a video, frame identifiers and time points can be marked for each obtained video frame. A frame identifier is used to uniquely mark a video frame, and different video frames can be distinguished based on the frame identifier. The time point refers to the time when the video frame appears in the video.
[0208] In some embodiments, during the process of the multi-modal understanding module 530 performing frame splitting on a video using computer vision services, classification labels for each video frame can be identified. The classification labels can be used to indicate the category to which the object shown in the video frame belongs. Exemplarily, the classification labels can be a person, a plant, or an animal shown in the video frame, etc. Another exemplarily, the classification labels can be a building, a natural scenery, etc. shown in the video frame.
[0209] It should be noted that the number of classification labels for video frames is not limited. For example, as Figure 3a shown, the classification labels of the thumbnail of "Video A" included in the first search result display interface 330 can include "boy", "star", "sky", and "tree", etc.
[0210] Step 2: The multi-modal understanding module 530 uses computer vision services to perform segmentation processing on the video.
[0211] Segmentation processing refers to using the video segmentation algorithm provided by computer vision services to divide the video into multiple video segments.
[0212] It can be understood that when a video is played, the content it shows changes continuously as the video frames are played in sequence, but the degree of change varies. Therefore, similar video frames (with a small degree of change) can be classified into one video segment.
[0213] In some embodiments, the degree of change between a video frame and its adjacent video frames can be measured based on the classification labels and image parameters of the video frames, and then the video can be segmented. Exemplarily, the image parameters of the video frame can include the jitter degree, clarity, pixel value, etc. of the video frame, and the present application does not limit this.
[0214] Among them, the jitter of a video frame refers to the phenomenon that the content shown in the video frame jitters or shakes during video playback. Exemplarily, when the user holds the mobile phone 300 to take a photo and moves the mobile phone 300 when hoping to shoot another scene, there may be an obvious jitter situation. The clarity of a video frame refers to the clarity of each detail texture and its boundary in the video frame. The pixels of a video frame can represent the brightness of the video frame.
[0215] It is understandable that when a video frame is compared with its adjacent video frames, the value corresponding to the jitter degree, the degree of change of other image parameters, and the degree of change of the classification label are positively correlated with the degree of change of the content presented by the video. That is, the larger or smaller the value corresponding to the jitter degree, the more obvious the degree of change of the image parameters and the classification label, and the more obvious the degree of change of the content presented by the video.
[0216] In some embodiments, the video segmentation algorithm can be represented by the following formula 1. Based on the classification label and image parameters of the video frame, the segmentation score of the video frame is calculated by combining the following formula 1, and then it is determined whether to determine it as the start frame or the end frame of a video segment according to the segmentation score of the video frame. Formula 1 is as follows:
[0217] y = α × frame A + β × frame B + γ × frame Y + δ × frame T
[0218] Wherein, y represents the segmentation score of the video frame, frame A represents the jitter degree score of the video frame, frame B represents the clarity change score of the video frame, frame Y represents the label change score of the video frame, frame T represents the pixel change score of the video frame, and α, β, γ, and δ respectively represent frame A 、frame B 、frame Y and frame T coefficients. In some embodiments, α, β, γ, and δ can be values preset manually according to the influence degree of the jitter degree, clarity change, label change, and pixel change of the video frame on the segmentation score of the video frame.
[0219] It should be noted that the above process of determining the segmentation score of the video frame based on the classification label and multiple image parameters is only an example. It can also be based on one or more of the classification label and multiple image parameters to determine the segmentation score of the video frame, and the present application does not limit this.
[0220] In some embodiments, video jitter detection methods such as the optical flow method, feature point matching method, and based on image gray distribution characteristics of image displacement can be used to detect the jitter degree of multiple video frames in the video. Then, based on the corresponding relationship between the preset jitter degree range and the jitter degree score, the jitter degree score of the video frame is determined.
[0221] In some embodiments, a clarity detection tool can be used to determine the clarity of multiple video frames in a video, and then the clarity of a video frame is compared with the clarity of adjacent video frames to determine the clarity change value of the video frame. Subsequently, based on the correspondence between a preset clarity change value range and clarity change scores, the clarity change score of the video frame is determined. Exemplarily, the video quality detection tool can be open-source software such as FFmpeg and Video Quality Measurement Tool.
[0222] In some embodiments, the classification label of a video frame can be compared with the classification labels of adjacent video frames, and the comprehensive label change of the video frame is determined based on the change in the number of classification labels of the video frame and the change in the content of the classification labels. Exemplarily, the union and intersection of the classification label of a video frame and the classification labels of its adjacent video frames can be calculated. The change in the number of classification labels in the union can reflect the change in the number of classification labels of the video frame, and the change in the number of classification labels in the intersection can reflect the change in the content of the classification labels. When there is a change in the number of classification labels in the union or the number of classification labels in the intersection, the label change score of the video frame can be determined according to the correspondence between a preset range of changes in the number of classification labels and label change scores.
[0223] In some embodiments, a pixel detection tool can be used to determine the pixels of multiple video frames in a video, and then the pixels of a video frame are compared with the pixels of adjacent video frames to determine the pixel change value of the video frame. Subsequently, based on the correspondence between a preset pixel change value range and pixel change scores, the pixel change score of the video frame is determined. Exemplarily, the pixel detection tool can be plugins such as PixelStick, MeasureIt, and Guides.
[0224] In some embodiments, the degree of change in the content presented by two adjacent video frames in a video can be positively correlated with the size of the segmentation score of the video frame, that is, the larger the segmentation score of the video frame, the more obvious the degree of change in the content presented compared with adjacent video frames.
[0225] Exemplarily, a video frame can be compared with its previous video frame, and a segmentation score threshold is preset, and the video frame whose segmentation score exceeds the segmentation score threshold is determined as the starting frame of a video segment. As shown in Formula 1 above, the higher the jitter of the video frame, the higher the corresponding value; the greater the clarity change of the video frame compared with the previous video frame, the higher the corresponding value; the greater the label change of the video frame compared with the previous video frame, the higher the corresponding value; the greater the label change of the video frame compared with the previous video frame, the higher the corresponding value; A the higher the corresponding value; the greater the clarity change of the video frame compared with the previous video frame, the higher the corresponding value; B the higher the corresponding value; the greater the label change of the video frame compared with the previous video frame, the higher the corresponding value; Y the higher the corresponding value; the greater the label change of the video frame compared with the previous video frame, the higher the corresponding value;Y The higher the corresponding value; the greater the pixel change of this video frame compared to the previous video frame, frame T The higher the corresponding value.
[0226] Exemplarily, a video frame can be compared with its next video frame, and a segmentation score threshold can be preset, and the video frame whose segmentation score exceeds the segmentation score threshold is determined as the end frame of a video segment.
[0227] In some embodiments, the degree of change in the content shown by two adjacent video frames in a video can also be negatively correlated with the magnitude of the segmentation score of the video frame, that is, the smaller the segmentation score of the video frame, the more obvious the change in the content shown compared to the adjacent video frame, that is, the lower the segmentation score of the video frame, the greater the possibility of determining it as the start frame or the end frame of a video segment.
[0228] Exemplarily, a video frame can be compared with its previous video frame, and a segmentation score threshold can be preset, and the video frame whose segmentation score is lower than the segmentation score threshold is determined as the start frame of a video segment. Combining the above formula 1, the higher the jitter degree of this video frame, frame A The lower the corresponding value; the greater the clarity change of this video frame compared to the previous video frame, frame B The lower the corresponding value; the greater the label change of this video frame compared to the previous video frame, frame Y The lower the corresponding value; the greater the label change of this video frame compared to the previous video frame, frame Y The lower the corresponding value; the greater the pixel change of this video frame compared to the previous video frame, frame T The lower the corresponding value.
[0229] Step 3: The multimodal understanding module 530 uses computer vision services to determine the representative frame corresponding to each video segment.
[0230] After obtaining multiple video segments, the multimodal understanding module 530 determines the representative frame that can represent the video semantics of this video segment from each video segment.
[0231] In some embodiments, the representative frame can be the start frame, end frame, middle frame or random frame in a video segment.
[0232] Exemplarily, assume that a video segment contains 99 video frames. Then, the 1st video frame (starting frame), the 99th video frame (ending frame), or the 50th video frame (middle frame) among these 99 video frames can be used as the representative frame of this video segment according to the time sequence. Alternatively, a video frame randomly selected from these 99 video frames (random frame) can be used as the representative frame of this video segment.
[0233] In some embodiments, the optimal frame can be determined from a video segment as the representative frame based on a preset rule related to the image parameters and classification labels of the video frames.
[0234] Exemplarily, the score of a video frame can be calculated based on the image parameters and classification labels of the video frame. The more labels a video frame has, the higher its score. The lower the jitter degree of a video frame, the higher its score. The higher the clarity of a video frame, the higher its score. The lower the pixel change of a video frame compared to the pixels of the previous video frame, the higher its score. Finally, the video frame with the highest score (i.e., the optimal frame) in a video segment can be determined as the representative frame.
[0235] Next, in combination with Figure 6 the process of determining the representative frame in the video will be introduced in detail.
[0236] As Figure 6 shown, this video includes 120 video frames. A video frame is compared with the previous video frame, and the segmentation scores corresponding to these 120 video frames are calculated in combination with the above formula 1. The video frames whose segmentation scores exceed the segmentation score threshold are used as the starting frames of a video segment.
[0237] In some embodiments, these 120 video frames are divided into 3 video segments, namely video segment 1, video segment 2, and video segment 3, and the middle frame of each video segment is determined as its corresponding representative frame. Video segment 1 includes 40 video frames, video segment 2 includes 51 video frames, and video segment 3 includes 29 video frames. That is, the 20th video frame in video segment 1 is used as the representative frame corresponding to video segment 1, the 26th video frame in video segment 2 is used as the representative frame corresponding to video segment 2, and the 15th video frame in video segment 3 is used as the representative frame corresponding to video segment 3.
[0238] It should be noted that the number of video frames in video segment 1 is 40, which is an even number. Both the 20th video frame and the 21st video frame are middle frames. Any one of the middle frames can be determined as the representative frame corresponding to video segment 1. This application does not make any limitations in this regard.
[0239] S504: The multimodal understanding module 530 returns the representative frame of the video and its related information to the gallery service module 510.
[0240] The relevant information of the representative frame includes, but is not limited to, the time point corresponding to the starting frame, the time point corresponding to the ending frame, and the time point corresponding to the representative frame in the video segment corresponding to the representative frame, the classification label of the video segment, etc. Among them, the time points corresponding to the starting frame, the ending frame, and the representative frame in the video segment are conducive to jumping to the corresponding video frame when displaying video search results for users later, improving the user experience. The classification label of the video segment is conducive to displaying more accurate video search results when performing video searches later.
[0241] In some embodiments, the classification label of the video segment can be used to indicate the type of the object shown in the video segment. Exemplarily, the union of the classification labels corresponding to the multiple video frames included in the video segment can be obtained to get the classification label of the video segment.
[0242] In addition, the relevant information of the representative frame can also include the name of the person and the relationship between the characters corresponding to the video segment. In some embodiments, after determining the representative frame, based on the preset names of the people and the relationship between the characters, the name of the person and the relationship between the characters corresponding to the person in the representative frame can be identified.
[0243] S505: The gallery service module 510 stores the representative frame of the video and its relevant information.
[0244] After receiving the representative frame and its relevant information returned by the multimodal understanding module 530, the gallery service module 510 can store them. In some embodiments, the gallery service module 510 is configured with a database, and the gallery service module 510 can store the representative frame of the video and its relevant information in the database.
[0245] S506: The gallery service module 510 calls the multimodal understanding module 530 to determine the visual semantic vector corresponding to the representative frame.
[0246] In some embodiments, the gallery service module 510 calls the computer vision service provided by the multimodal understanding module 530 to determine the visual semantic vector corresponding to the representative frame. A video includes multiple video segments, and each video segment corresponds to a representative frame, so a video corresponds to multiple visual semantic vectors. The gallery service module 510 calls the computer vision service provided by the multimodal understanding module 530 to perform video semantic understanding on the video, and further determines the visual semantic vectors corresponding to the multiple representative frames in the video.
[0247] In some embodiments, the computer vision service can provide a CLIP model. The representative frame can be input into the image encoder of the CLIP model to obtain the visual semantic vector corresponding to the representative frame.
[0248] In some embodiments, the image encoder of the CLIP model can be trained with image training samples, which may include picture training samples and video frame training samples. The training process of the image encoder of the CLIP model can be referred to in the following embodiments.
[0249] S507: The multimodal understanding module 530 returns the visual semantic vector corresponding to the representative frame to the gallery service module 510.
[0250] S508: The gallery service module 510 stores the visual semantic vector corresponding to the representative frame.
[0251] Upon receiving the visual semantic vector corresponding to the representative frame returned by the multimodal understanding module 530, the gallery service module 510 can store it. Exemplarily, if the gallery service module 510 is configured with a database, the gallery service module 510 can store the visual semantic vector corresponding to the representative frame in the database.
[0252] S509: The gallery service module 510 sends the representative frame of the video, its related information, and the visual semantic vector of the representative frame to the search module 520.
[0253] S510: The search module 520 constructs an index corresponding to the video.
[0254] An index is an index corresponding to a video segment in the video. A video includes multiple video segments, so a video corresponds to multiple indexes, and different indexes correspond to different video segments in the video.
[0255] The search module 520 can combine the visual semantic vector of the representative frame, the related information of the representative frame, and the attribute tags of the video segment to construct an index corresponding to the video segment of the representative frame. As exemplified above, the index of the video segment may include the visual semantic vector of the representative frame of the video segment, the time points corresponding to the start frame, the end frame, and the representative frame in the video segment, the classification label of the video segment, and the attribute tags of the video segment (such as the video acquisition time, the video acquisition location, and the video storage path), etc.
[0256] It should be noted that the content included in the index of the above video segment is only an example, and the index of the video segment may include any information among the visual semantic vector of the representative frame, the related information of the representative frame, and the attribute tags of the video segment. This application does not make any limitations in this regard.
[0257] Exemplarily, the search module 520 may include an index library, and the search module 520 can store the index corresponding to the video in the index library for subsequent search and matching based on the index library.
[0258] It can be understood that in practical applications, the number of videos stored in the mobile phone 300 may be very large, and a video corresponds to multiple indexes, indicating that the index library may contain a large number of indexes. In some embodiments, in order to improve the search efficiency, an inverted index method can be used for searching.
[0259] Exemplarily, after the search module 520 constructs the indexes corresponding to the video, the visual semantic vectors included in the indexes in the index library are subjected to vector clustering to obtain a plurality of clustering clusters, that is, the vector space corresponding to all indexes is divided into a plurality of vector regions, and a clustering cluster is included in one vector region. Each vector region includes a plurality of indexes with relatively high vector similarity, and each vector region can be replaced by a clustering center point. For example, methods such as K-means or hierarchical clustering can be used for clustering, and the present application does not limit this.
[0260] In this way, when the mobile phone 300 performs a search match based on the text semantic vector of the search text, it can first match with the clustering center point, and then match with the visual semantic vectors in the vector region to which the determined clustering center point belongs, without having to match with all the indexes in the index library, which not only saves computing resources but also greatly reduces the time consumed by the search match, can avoid search latency, improve the search efficiency of the video, and further enhance the user experience.
[0261] In some embodiments, the search module 520 can store the received representative frame and its related information and the visual semantic vector of the representative frame, which is convenient for searching when constructing indexes, thereby improving the construction speed of the indexes.
[0262] Exemplarily, the visual semantic vector of the representative frame corresponding to a video segment and the related information of the representative frame can be stored in a document. Taking video segment 1 as an example, the visual semantic vector of the representative frame corresponding to video segment 1 and the related information of the representative frame can be stored in sub-document 1. The information stored in sub-document 1 can be seen in Table 2 below:
[0263] Table 2
[0264] Information Name Information content segments.media-vector [bd de d1 b4 3c 8c 9c] segments.startTime 0 segments.endTime 91666 segments.startFrame 0 segments.endFrame 2750 segments.tag-name People|Landscape|Architecture
[0265] As shown in Table 2, segments.media-vector represents the visual semantic vector of the representative frame in video segment 1. In practical applications, a vector is usually composed of an array. In the embodiments of the present application, a hexadecimal sequence process is performed on the visual semantic vector to obtain a visual semantic vector in the form of [bd de d1 b4 3c 8c 9c]. The hexadecimal floating-point number enables the mobile phone 300 to store it with less storage space, which can save storage space. segments.startTime represents the start time of video segment 1, which is 0 ms (that is, the time point corresponding to the start frame in video segment 1). segments.endTime represents the end time of video segment 1, which is 91666 ms (that is, the time point corresponding to the end frame in video segment 1). segments.startFrame and segments.endFrame represent that there are 2750 video frames in video segment 1 from start to end. segments.tag-name represents the classification tags of video segment 1, including "person", "scenery", and "building".
[0266] In some embodiments, the attribute tags of video segment 1 can also be stored in the sub-document 1, and the attribute tags of video segment 1 are the attribute tags of the video stored in step S502. Taking video segment 1 as an example of a video segment in the video captured by the user, the attribute tags of video segment 1 can include the shooting time, shooting location, and storage path in the mobile phone 300, etc.
[0267] In some embodiments, the attribute tags of video segment 1 can also be stored in document 1. It can be understood that the attribute tags of a video are fixed, that is, the attribute tags corresponding to multiple video segments in a video are the same. Then the attribute tags of a video can also be stored in document 1, and the attribute tags of each video segment of the video can be obtained from document 1, which can reduce the content pressure of the mobile phone 300 and reduce the storage cost.
[0268] Exemplarily, video D includes the above-mentioned video segment 1, and also includes video segment 2 and video segment 3. The visual semantic vector of the representative frame corresponding to video segment 1 and the relevant information of the representative frame are stored in sub-document 1. The visual semantic vector of the representative frame corresponding to video segment 2 and the relevant information of the representative frame can be stored in sub-document 2. The visual semantic vector of the representative frame corresponding to video segment 1 and the relevant information of the representative frame can be stored in sub-document 3. Then the information stored in document 1 can be seen as shown in Table 3 below:
[0269] Table 3
[0270]
[0271]
[0272] As shown in Table 3, file_path represents the storage path of video D in the mobile phone 300. shooting-time represents the shooting time of video D. location represents the shooting location of video D. segments is used to indicate the sub-documents corresponding to multiple video segments included in video D. Based on Table 2 and Table 3 above, the attribute tags of video segment 1 can be obtained from document 1, and the relevant information of the representative frame of video segment 1 can be obtained from sub-document 1. Similarly, the attribute tags of video segment 2 and video segment 3 can also be obtained from document 1.
[0273] It should be noted that since video semantic understanding may consume a large amount of computing resources, in order to avoid affecting the user's use, such as causing lags, steps S503 - S510 can be executed when the mobile phone 300 is in the state of being charged and the screen is off.
[0274] The video search method provided by this application searches and matches based on the search text input by the user and the index corresponding to the video. The index corresponding to the video is the index of a video segment in the video, and the index of the video segment includes at least the visual semantic vector of the representative frame in the video segment. The visual semantic vector of the representative frame can indicate the meaning expressed by the content of the picture shown by the representative frame, and it can represent the video semantics of the video segment. Therefore, the visual semantic vector of the representative frame is fully associated with the video semantics of the video segment in the video. Searching and matching based on the visual semantic vector of the representative frame can achieve the fusion interaction between the search text and the content of the picture shown by the representative frame, thereby improving the accuracy of the video search results and enhancing the user's experience.
[0275] Next, in combination with Figure 5b Continue to introduce in detail the steps included in the search stage of the video search method provided by this application.
[0276] S511: The gallery service module 510 receives the input operation of the user for the search text.
[0277] The user can input the search text in the search interface provided by the mobile phone 300. Exemplarily, the user can input the search text in the search box 321 included in the album display interface 320 as shown in Figure 3a Exemplarily, the user can input the search text in the search box 421 included in the negative first screen interface 420 as shown in Figure 4a Exemplarily, the user can input the search text in the search box 421 included in the negative first screen interface 420 as shown in
[0278] The search text is the text in which the user describes the characteristics of the video they need. Exemplarily, the search text may include the video acquisition time, the video acquisition location, and the content of the video screen, etc. For example, the search text may be "scenery shot last week", and the present application does not limit this.
[0279] S512: The gallery service module 510 sends the search text to the search module 520.
[0280] S513: The search module 520 calls the multimodal understanding module 530 to determine the text semantic vector corresponding to the search text.
[0281] The search module 520 calls the multimodal understanding module 530 to perform text semantic understanding on the search text, and obtains the text semantic vector corresponding to the search text.
[0282] Text semantic understanding refers to enabling the mobile phone to understand the meaning expressed by the text, and it is a key technology of natural language processing (NLP) technology.
[0283] In some embodiments, the multimodal understanding module 530 provides a CLIP model, and the search text can be input into the text encoder of the CLIP model to obtain the text semantic vector corresponding to the search text.
[0284] In some embodiments, the text encoder of the CLIP model can be trained through text training samples. The training process of the text encoder of the CLIP model can be referred to in the following embodiments.
[0285] S514: The multimodal understanding module 530 returns the text semantic vector corresponding to the search text to the search module 520.
[0286] S515: The search module 520 performs vector recall in the index library based on the text semantic vector corresponding to the search text.
[0287] Vector recall refers to recalling the index that matches the text semantic vector corresponding to the search text in the index library.
[0288] In some embodiments, the vector similarity between the text semantic vector corresponding to the search text and the visual semantic vectors respectively included in multiple indexes in the index library can be calculated, and the vector similarity calculation results respectively corresponding to the multiple indexes are obtained. And N indexes with higher vector similarity among the multiple vector similarity calculation results are used as the vector recall results. Among them, N is an integer greater than 0. Exemplarily, N can be the number of vector recall results set in advance, such as 5, 8 or 10, etc.
[0289] Vector similarity refers to the degree of similarity between two vectors, which can be calculated using various methods. Exemplarily, the similarity degree can be determined by calculating the cosine similarity between two vectors, or other methods can be used. This application does not make any limitations in this regard.
[0290] In some embodiments, as described above, the vector space in the index library includes multiple vector regions, and each vector region includes multiple indexes with relatively high vector similarity. Then, in the embodiments of this application, the distance between the text semantic vector corresponding to the search text and the clustering center points of multiple vector regions in the index library can be calculated to determine the clustering center point with the closest distance. Then, calculate the vector similarity between the visual semantic vectors included in the multiple indexes in the vector region to which the clustering center point belongs and the text semantic vector. Sort the indexes in descending order of vector similarity to obtain the inverted index chain corresponding to the clustering center point. Exemplarily, the first N indexes in the inverted index chain can be used as the vector recall result. Another example is that the indexes whose vector similarity exceeds the vector similarity threshold can be used as the vector recall result.
[0291] The following Figure 7 Specifically illustrate the process of vector recall.
[0292] Such as Figure 7 As shown, the inverted index library includes multiple clustering center points such as clustering center point 1. First, calculate the distance between the text semantic vector corresponding to the search text and the clustering center points of multiple vector regions in the index library, and determine that clustering center point 1 is the clustering center point with the closest distance. Then, calculate the vector similarity between the visual semantic vectors of the multiple indexes in the vector region to which center 1 belongs and the text semantic vector, and sort them in descending order of vector similarity to obtain the inverted index chain 1 corresponding to clustering center point 1. Select the TopN indexes from the inverted index chain 1 corresponding to clustering center point 1 as the vector recall result.
[0293] Among them, in the inverted index chain 1, the visual semantic vector corresponding to index 1 is the closest to clustering center point 1, and the distances between index 2 and index 3 and clustering center point 1 gradually become farther. Therefore, the selected TopN indexes are sequentially selected backward starting from index 1. N can be any integer greater than 0, and this application does not make any limitations in this regard.
[0294] S516: The search module 520 calls the natural language understanding module 540 to identify the entities in the search text.
[0295] The search module 520 calls the natural language understanding module 540 to identify the entities included in the search text.
[0296] Exemplarily, entities with specific meanings in the search text can be recognized through Named Entity Recognition (NER). Entities can include, but are not limited to, time, location, person names, organization names, and proper nouns. Taking the search text "scenery taken last week" as an example, the entities in this search text include: "last week" and "scenery".
[0297] S517: The natural language understanding module 540 returns the entities in the search text to the search module 520.
[0298] S518: The search module 520 performs entity recall in the index library based on the entities in the search text.
[0299] Entity recall refers to recalling the indexes in the index library that match the entities in the search text.
[0300] In some embodiments, the index may include relevant information representing frames in the video segment and attribute tags of the video segment, and entities are included therein. For example, the video acquisition time, video acquisition location, classification tags of the video segment, etc. may all include entities. Taking the search text "scenery taken in City B last week" as an example, there are entities "City B" as the location, "last week" as the time, and the entity "scenery" related to the content of the video display screen. Then, matching can be performed among the entities corresponding to multiple indexes to obtain the indexes that match the entities in the search text as the entity recall results.
[0301] S519: The search module 520 sorts the vector recall results and entity recall results.
[0302] In some embodiments, the intersection result or union result of the vector recall result and the entity recall result can be sorted.
[0303] Exemplarily, sorting can be performed according to the vector similarity between the vector recall result and the search text, and the entity matching degree between the entity recall result and the search text. For example, the vector similarity between the text semantic vector of the search text and the visual semantic vector of the recall result (vector recall result or entity recall result), and the matching degree between the entities in the search text and the entities in the recall result can be weighted and summed to obtain the comprehensive matching degree corresponding to the recall result, and the recall results (including vector recall results and entity recall results) are sorted in descending order according to the corresponding comprehensive matching degrees.
[0304] Thus, based on the search text input by the user, on the basis of performing a search and match of the text semantic vector, a search and match of the entities in the search text is performed, and the final display result order is obtained based on the comprehensive match degree of the search results, ensuring that the videos presented to the user are more matching results with the user's search text, and further improving the user experience.
[0305] S520: The search module 520 returns the search results to the gallery service module 510.
[0306] The search module 520 returns the sorted search results to the gallery service module 510.
[0307] S521: The gallery service module 510 presents the search results to the user.
[0308] It can be understood that in practical applications, the mobile phone 300 usually stores pictures and videos in the gallery application. Therefore, when searching in the search interface provided in the gallery application, both picture search results and video search results will be displayed. That is, the index library includes not only the indexes corresponding to the video segments, but also the indexes corresponding to the pictures.
[0309] In some embodiments, the index corresponding to the picture may include, but is not limited to, the picture semantic vector, the attribute tags of the picture, etc. Exemplarily, as described above, a video frame can be regarded as an image, and a picture is also an image. Therefore, the picture semantic vector corresponding to the picture can be generated by the image encoder of the CLIP model provided by the multi-modal understanding module 530 and returned to the gallery service module 510. Exemplarily, the attribute tags of the picture can be obtained by the gallery service module 510 first receiving and storing the user's new operation or modification operation on the attribute tags of the picture. The search module 520 then receives the picture semantic vector and the attribute tags of the picture sent by the gallery service module 510 and constructs the index corresponding to the picture based on this.
[0310] Exemplarily, as Figure 3a of the first search result display interface 330, the display shows that the first search results include video search results and picture search results.
[0311] In some embodiments, the search results can be presented based on multiple sorted indexes. As described above, the index corresponding to the video includes the visual semantic vector representing the frame, the relevant information representing the frame, and the attribute tags of the video segment. The index corresponding to the picture includes the picture semantic vector and the attribute tags of the picture, etc.
[0312] Exemplarily, as Figure 3aThe first search result display interface 330, the picture search results in the first display area 331 show images, and the video search results show thumbnails and time points. The information corresponding to the thumbnails and time points is the information included in the index corresponding to the video search results. Taking Figure 3a "Video A" in
[0313] as an example, its thumbnail is the starting frame of video segment a included in the index, its first time point is the time point "02:18" corresponding to the starting frame of video segment a included in the index, and its second time point is the total duration of the video "08:32" included in the index. Figure 3b As shown in (2) of
[0314] the picture search results include the shooting time of the picture "October 1, 2023" and the shooting location of the picture "City B".
[0315] It should be noted that the gallery service module, the search module, the multimodal understanding module, and the natural language understanding module can also be located in the cloud server. That is, the cloud server uses the interaction of these four modules to implement the steps included in the index construction stage. In the search stage, it can be based on the interaction between an electronic device such as the mobile phone 300 and the cloud server to implement the steps included in the search stage.
[0316] In some embodiments, the mobile phone 300 can send the search text input by the user to the gallery service module of the cloud server, so that the gallery service module of the cloud server interacts with other modules to implement the steps included in the search stage. Then, the gallery service module of the cloud server sends the search results to the mobile phone 300, so that the mobile phone 300 can display the search results to the user.
[0317] In some embodiments, it is assumed that the total duration of video E is 2 minutes and it is shot at 30 frames per second, that is, 30 images are shot in one second. When video E is stored in the cloud server, the multimodal understanding module in the cloud server can perform frame splitting on video E, decompose video E into individual video frames, and 3600 video frames can be obtained; the multimodal understanding module then performs segmentation on video E, segments it in units of 1 second, and divides it into 120 video segments. Each video segment includes 30 video frames (that is, 30 images shot in one second); the multimodal understanding module scores the 30 video frames in each video segment, and takes the video frame with the highest score as the representative frame of the video segment. Exemplarily, the scoring can be based on the jitter, clarity, pixels, etc. of the video frame, and the present application does not limit this.
[0318] Subsequently, the search module of the cloud server uses the visual semantic vector of the representative frame as the index corresponding to each video segment of video E, so that subsequent matching can be performed based on the search text and the representative frame of each video segment of video E. If the search text matches the 50th video segment of video E successfully, the thumbnail of the video returned to the user is the representative frame of the 50th video segment, and the returned time point is 50s. For the specific implementation manners of the above embodiments, reference can be made to Figure 5b the introduction in the above embodiments and will not be elaborated here.
[0319] It should be noted that the multimodal understanding module 530 of the mobile phone 300 can also segment the videos stored in the picture library in units of 1 second, and the present application does not limit this.
[0320] In the above embodiments, the video search method provided by the present application needs to apply the image encoder of the CLIP model and the text encoder of the CLIP model, so the image encoder of the CLIP model and the text encoder of the CLIP model need to be trained first. In some embodiments, the image encoder of the CLIP model and the text encoder of the CLIP model can be trained separately. In some embodiments, the image encoder of the CLIP model and the text encoder of the CLIP model can be jointly trained based on the training method of contrastive learning.
[0321] The following combines Figure 8 and Fig. 9 to introduce the training process of the image encoder of the CLIP model and the text encoder of the CLIP model. The following embodiments will introduce the training method of contrastive learning of the image encoder of the CLIP model and the text encoder of the CLIP model in detail in the following steps 1-step 7.
[0322] Step 1: Obtain image training samples and text training samples corresponding to the image training samples.
[0323] Both the image encoder of the CLIP model and the text encoder of the CLIP model need to be pre-trained with a large number of training samples. Therefore, before model training, it is necessary to obtain the training samples of the image encoder of the CLIP model, that is, image training samples, and obtain the training samples of the text encoder of the CLIP model, that is, the text training samples corresponding to the image training samples.
[0324] As described above, the image training samples can include picture training samples and video frame training samples. The picture training sample can be any picture, and the video frame training sample can be a video frame in any video. The text training sample corresponding to the image training sample refers to the text corresponding to the content shown in the image training sample, that is, the text training sample can express the content shown in the image training sample. Exemplarily, if the image training sample is Figure 3a the thumbnail shown by "Video A" in the first display area 331 in
[0325] This application does not limit the method for obtaining the text training samples corresponding to the image training samples.
[0326] Exemplarily, the text training samples corresponding to the image training samples can be manually labeled, that is, manually labeled according to one's own understanding of the image semantics of the image training samples. Also exemplarily, the text training samples corresponding to the image training samples can be automatically generated by recognizing relevant content such as objects, scenes, and actions in the image training samples. Additionally exemplarily, the text training samples corresponding to the image training samples can be automatically generated by a text generation model for generating descriptive texts of images.
[0327] It should be noted that this application does not limit the number of image training samples. It can be understood that the text training samples correspond to the image training samples, so the numbers of the two are the same.
[0328] As Fig. 9 shown, obtain N image training samples, and obtain N text training samples corresponding one-to-one to the N image training samples. Exemplarily, image training sample 1 corresponds to text training sample 1.
[0329] Step 2: Input the image training samples into the image encoder, and the image encoder outputs the image vectors corresponding to the image training samples.
[0330] As Figure 8 shown, for the image training samples, the image encoder can encode them to obtain the image vectors of the image training samples.
[0331] As Fig. 9As shown, input N image training samples into the image encoder to obtain the image vectors I corresponding to the N image training samples respectively 1 、I 2 、I 3 ……I N 。
[0332] Step 3: Input the text training samples into the text encoder, and the text encoder outputs the text vectors corresponding to the text training samples
[0333] As Figure 8 shown, for the text training samples, the text encoder can encode them to obtain the text vectors of the text training samples
[0334] As Fig. 9 shown, input N text training samples into the text encoder to obtain the text vectors T corresponding to the N text training samples respectively 1 、T 2 、T 3 ……T N 。
[0335] Step 4: Combine each image vector with multiple text vectors respectively to obtain multiple vector pairs, and determine the vector pairs with corresponding relationships among the multiple vector pairs as positive sample vector pairs, and determine the remaining vector pairs as negative sample vector pairs
[0336] Contrastive learning is an unsupervised training method, so it is necessary to define positive and negative samples from the training samples. In the embodiments of the present application, it is to determine positive sample vector pairs and negative sample vector pairs from multiple vector pairs
[0337] In some embodiments, assuming there are N image vectors and N text vectors, combine each image vector with the N text vectors respectively to obtain N×N vector pairs. It can be understood that the image training samples and the text training samples have corresponding relationships, so the vectors corresponding to them also have corresponding relationships. Among the N×N vector pairs, the vector pairs composed of the image vectors and text vectors with corresponding relationships are determined as positive sample vector pairs, that is, there can be N positive sample vector pairs; the remaining vector pairs are determined as negative sample vector pairs, that is, there can be N×(N - 1) negative sample vector pairs
[0338] As Fig. 9 shown, taking I 1 as an example, it is combined with T 1 、T 2 、T 3 ……T N respectively to obtain I 1 ˙T 1 、I 1 ˙T2 and I 1 ˙T 3 ……I 1 ˙T N These N vector pairs, I 2 and I 3 ……I N Similarly, N×N vector pairs can be obtained. Take 1 ˙T 1 and I 2 ˙T 2 and I 3 ˙T 3 ……I N ˙T N The vector pairs with such a corresponding relationship are determined as positive sample vector pairs, and the rest are determined as negative sample vector pairs.
[0339] Step 5: Calculate the vector similarity between the image vector and the text vector in each vector pair.
[0340] Exemplarily, the vector cosine similarity between the image vector and the text vector in each vector pair can be calculated.
[0341] Step 6: Based on the loss function, the vector similarity corresponding to the positive sample vector pair, and the vector similarity corresponding to the negative sample vector pair, adjust the parameters of the image encoder and the text encoder.
[0342] As Figure 8 shown, contrastive learning is performed based on the image encoder and the image vector output by it, as well as the text encoder and the text vector, to adjust the parameters of the image encoder and the text encoder.
[0343] It can be understood that in the embodiments of the present application, the model training objective is to maximize the vector similarity corresponding to the positive sample vector pair and minimize the vector similarity corresponding to the negative sample vector pair.
[0344] Exemplarily, the loss function can be a cross-entropy loss function, and its formula is specifically as follows:
[0345]
[0346] Among them, represents the loss value of the loss function, K represents the number of vector pairs, y i represents the true vector similarity corresponding to the i-th vector pair, represents the predicted vector similarity corresponding to the i-th vector pair.
[0347] In addition, in some embodiments, the true vector similarity corresponding to the positive sample vector pair can also be represented as 1, and the true vector similarity corresponding to the negative sample vector pair can be represented as 0. The parameters of the image encoder and the text encoder are adjusted until the predicted vector similarity corresponding to the positive sample vector pair can approach 1 to the greatest extent, and the predicted vector similarity corresponding to the negative sample vector pair can approach 0 to the greatest extent, that is, the value of the loss function is minimized.
[0348] Step 7: When the training termination condition is satisfied, end the training to obtain a trained image encoder and a trained text encoder. Exemplarily, the training termination condition can be that the preset number of training times is reached during the model training process, or the loss value of the loss function is less than the loss value threshold during the model training process, etc.
[0349] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a computer, it can implement one or more steps in any of the above video search methods.
[0350] The computer-readable storage medium can be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium can be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage devices, etc.
[0351] Another embodiment of the present application also provides a computer program product containing instructions. When the computer program product is executed by a computer, it can implement one or more steps in any of the above video search methods.
[0352] The electronic device, computer-readable storage medium, and computer program product provided in this embodiment are all used to execute the corresponding video search method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding video search method provided above, and will not be elaborated here.
[0353] The terms "first", "second", "third", etc. in the specification, claims, and drawings of the present application are used to distinguish different objects, rather than to limit a specific order.
[0354] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
Claims
1. A video search method, characterized in that, the method includes: displaying a first interface of a gallery application; the first interface includes an input box, and the input box includes a first text; wherein, the first interface includes a first thumbnail of a first video, and the first thumbnail includes a first time point; in response to a triggering operation of the user on the first thumbnail, playing the first video starting from the first time point; displaying a second interface of the gallery application; the second interface includes the input box, and the input box includes a second text; wherein, the second interface includes a second thumbnail of the first video, and the second thumbnail includes a second time point; in response to the triggering operation of the user on the second thumbnail, playing the first video starting from the second time point; the second time point is later than the first time point.
2. The method according to claim 1, characterized in that, the first text is different from the second text, the first video matches the first text, and the first video matches the second text.
3. The method according to claim 2, characterized in that, the first video includes a first video frame and a second video frame; the first video matches the first text and the first video matches the second text, including: the first text matches the first video frame corresponding to the first thumbnail, the second text matches the second video frame corresponding to the second thumbnail, and the second video frame is after the first video frame.
4. The method according to claim 3, characterized in that, the step of in response to a triggering operation of the user on the first thumbnail and playing the first video starting from the first time point includes: playing the first video starting from the first video frame or a first starting frame; the step of in response to the triggering operation of the user on the second thumbnail and playing the first video starting from the second time point includes: playing the first video starting from the second video frame or a second starting frame; wherein, the first starting frame is before the first video frame, the second starting frame is before the second video frame, and the second starting frame is after the first video frame.
5. The method according to any one of claims 1-4, characterized in that, the method further includes: displaying a third interface of the gallery application; the third interface includes the input box, the input box includes a third text, the third text is different from the first text, and the third interface includes the first thumbnail; in response to the triggering operation of the user on the first thumbnail, playing the first video starting from the first time point.
6. The method according to any one of claims 1-4, characterized in that, the first interface further includes a third thumbnail of the first video, and the third thumbnail includes a third time point; the method further includes: in response to the triggering operation of the user on the third thumbnail, playing the first video starting from the third time point; the first text matches the third video frame corresponding to the third thumbnail; the third time point is later than the second time point.
7. The method according to claim 6, wherein, said playing the first video starting from the third time point includes: playing the first video starting from the third video frame or the third starting frame; the third starting frame is before the third video frame and after the second video frame.
8. The method according to any one of claims 1-7, wherein, the first interface further includes a fourth thumbnail of a second video, the fourth thumbnail includes a fourth time point, and the method further includes: in response to a triggering operation of the user on the fourth thumbnail, playing the second video starting from the fourth time point; the first text matches a fourth video frame corresponding to the fourth thumbnail, and the second video is different from the first video.
9. A video search method, wherein, the method includes: displaying a first interface of a gallery application; the first interface includes an input box, and the input box includes a first text; wherein, the first interface includes a first thumbnail of a first video, the first thumbnail includes a first time point, and the first text is matched with a first video frame corresponding to the first thumbnail through a CLIP model; in response to a triggering operation of the user on the first thumbnail, playing the first video starting from the first time point; displaying a second interface of the gallery application; the second interface includes the input box, the input box includes a second text, and the second text is matched with a second video frame corresponding to the second thumbnail through the CLIP model; wherein, the second interface includes a second thumbnail of the first video, and the second thumbnail includes a second time point; in response to a triggering operation of the user on the second thumbnail, playing the first video starting from the second time point; the second time point is later than the first time point.
10. The method according to claim 9, wherein, the first video frame is in a first video segment of the first video, the second video frame is in a second video segment of the first video, and the first video frame and the second video frame are determined through the following steps: performing a first process on the first video to obtain a plurality of video frames and classification labels of each video frame; performing a second process on the first video based on the classification labels respectively corresponding to the plurality of video frames to obtain a plurality of video segments including the first video segment and the second video segment; determining a video frame of the first video segment as the first video frame, and determining a video frame of the second video segment as the second video frame.
11. The method according to claim 10, wherein, said performing a second process on the first video based on the classification labels respectively corresponding to the plurality of video frames to obtain a plurality of video segments including the first video segment and the second video segment includes: Performing a second process on the first video based on the differences between the image parameters respectively corresponding to adjacent video frames of the multiple video frames, and the differences between the classification labels respectively corresponding to adjacent video frames of the multiple video frames, to obtain multiple video segments including the first video segment and the second video segment.
12. The method according to claim 10, wherein, determining a video frame of the first video segment as the first video frame includes: determining the first starting frame, the first ending frame, a random video frame, or the video frame corresponding to the time point at the first position of the first video segment as the first video frame; the time point at the first position is determined based on the average of the starting time point and the ending time point of the first video segment.
13. The method according to claim 10, wherein, determining a video frame of the first video segment as the first video frame includes: determining the first video frame based on the image parameters respectively corresponding to the multiple video frames of the first video segment, and the classification labels respectively corresponding to the multiple video frames of the first video segment.
14. The method according to claim 9, wherein, the matching of the first text and the first video frame corresponding to the first thumbnail through the CLIP model includes the following steps: inputting the first text into the text encoder of the CLIP model to obtain the text semantic vector of the first text; inputting the first video frame into the image encoder of the CLIP model to obtain the visual semantic vector of the first video frame; matching the first text and the first video frame based on the text semantic vector and the visual semantic vector.
15. The method according to claim 14, wherein, the first vector similarity between the text semantic vector and the visual semantic vector is greater than or equal to a first threshold.
16. The method according to claim 14, wherein, the second vector similarity between the text semantic vector and the vector of the first clustering center point of the inverted index library is greater than or equal to a second threshold; the visual semantic vector of the first video frame belongs to the first clustering cluster corresponding to the first clustering center point; the third vector similarity between the text semantic vector and the visual semantic vector is greater than or equal to a third threshold; the inverted index library includes clustering clusters respectively corresponding to multiple clustering center points; the multiple clustering center points are determined by clustering multiple visual semantic vectors in the inverted index library.
17. The method according to any one of claims 10-16, wherein, the entity of the first text matches the entity of the attribute label of the first video segment.
18. The method according to any one of claims 9-16, wherein, the first interface further includes a third thumbnail of a second video, and the third video frame corresponding to the third thumbnail is in the third video segment of the second video; the entity of the first text matches the entity of the attribute label of the third video segment.
19. The method according to claim 17, wherein, The first interface further includes a fourth thumbnail of the first video. The first text matches the fourth video frame corresponding to the fourth thumbnail. The fourth video frame is in the fourth video segment of the first video. The entity of the first text matches the entity of the attribute label of the fourth video segment. The display order of the first thumbnail is before that of the fourth thumbnail. The display order of the first thumbnail and the fourth thumbnail is determined through the following steps: Based on the first vector similarity between the visual semantic vector of the first video frame and the text semantic vector of the first text, and the first matching degree between the attribute label of the first video segment and the entity of the first text, determine the first comprehensive matching degree between the first thumbnail and the first text. Based on the second vector similarity between the visual semantic vector of the fourth video frame and the text semantic vector of the first text, and the second matching degree between the attribute label of the fourth video segment and the entity of the first text, determine the second comprehensive matching degree between the second thumbnail and the first text. According to the order from largest to smallest of the comprehensive matching degrees, display the first thumbnail before the fourth thumbnail.
20. An electronic device Characterized in that it includes a memory, a display screen and a processor; The memory is coupled to the processor. The memory is used to store computer program code. The computer program code includes computer instructions. The display screen provides a display function. The one or more processors call the computer instructions to enable the electronic device to execute the video search method according to any one of claims 1-8, or the video search method according to any one of 9-19.
21. A computer-readable storage medium Characterized in that a computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the video search method according to any one of claims 1-8, or the video search method according to any one of 9-19.
Citation Information
Cited By
Visual media search method, electronic device and storage medium
EP4752699A1
Visual media search method, electronic device and storage medium
WO2025103081A1