Visual media searching method and electronic equipment

By implementing the visual media search method on an electronic device, search results can be sorted and displayed according to the degree of matching, the problem of poor image search experience in the prior art is solved, and the user experience and search efficiency are improved.

CN120067372APending Publication Date: 2025-05-30HONOR DEVICE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311580640.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, the image search function cannot effectively improve the user experience, especially when searching for videos, it is difficult for users to find content that meets the search intention.

Method used

A visual media search method is provided, performed by an electronic device, able to search videos and images, and sort the results according to the degree of matching with the search content, displaying the content that is most relevant to the user.

Benefits of technology

Improve the user experience. Through a consistent interface, users can find content that meets the search intention more quickly, enhancing the search precision and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067372A_ABST
    Figure CN120067372A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a visual media searching method and electronic equipment. The method comprises the steps that a first interface of a gallery application is displayed; the first interface comprises a search box; receiving a first operation of inputting a first text by a user through the search box; in response to the first operation, searching to obtain a plurality of hit results matched with the first text, the plurality of hit results including at least one hit image and at least one hit video segment, and the hit video segment being a video segment in the target video; respectively determining a matching score of each hit result, wherein the matching score represents the matching degree of the hit result and the first text; sorting the at least one hit image and the at least one target video according to the matching score of each hit result to obtain a first sorting result; and displaying the at least one hit image and the at least one target video according to the first sorting result. According to the method, image and video searching can be achieved, the searching results can be sorted, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic technologies, and particularly to a visual media search method and an electronic device. Background Art

[0002] With the rapid development of the camera configurations of electronic devices such as mobile phones and tablet computers, taking pictures and shooting videos with electronic devices have become one of the important ways for people to record their lives. Moreover, with the rapid development of aspects such as the storage capacity of electronic devices and network technologies, the images and videos stored in users' electronic devices are also increasing.

[0003] To facilitate users to manage and view the images in electronic devices, an image search function is configured in the gallery application or other similar applications of the electronic device. For example: The gallery application in the electronic device can classify the images in the electronic device according to information such as the time and location of image shooting, and users can view relevant images by searching for information such as time and location.

[0004] However, in related technologies, the experience brought by image search to users still needs to be improved. Summary of the Invention

[0005] This application provides a visual media search method and an electronic device, which can search for videos and images, and can sort the videos and images according to the matching degree with the search content, thereby improving the user experience.

[0006] In a first aspect, this application provides a visual media search method, which is executed by an electronic device. The electronic device can be a device including a gallery application such as a mobile phone, a tablet computer, or a laptop computer. The method includes: The electronic device displays a first interface of the gallery application; the first interface includes a search box; the electronic device receives a first operation in which the user inputs a first text through the search box; in response to the first operation, the electronic device searches for a plurality of hit results that match the first text, and the plurality of hit results include at least one hit image and at least one hit video segment, and at least one hit video segment is a video segment in at least one target video; the electronic device respectively determines the matching scores of the respective hit results, and the matching score represents the matching degree between the hit result and the first text; the electronic device sorts at least one hit image and at least one target video according to the matching scores of the respective hit results to obtain a first sorting result; the electronic device displays at least one hit image and at least one target video according to the first sorting result.

[0007] The first interface can be, for example, an album interface, a photo interface, an interface of the moment function, or a creation interface of the gallery application. The first text is the search text input by the user, such as "a girl doing yoga", "a boy standing near a bridge", etc.

[0008] In the embodiments of the present application, the hit results include both images (referred to as hit images) and video segments (referred to as hit video segments). The electronic device separately scores the matching degrees of the hit images and the hit video segments to obtain matching scores. Then, based on the matching scores, the videos to which the hit images and the hit video segments belong (referred to as target videos, also referred to as hit videos) are sorted. Optionally, the higher the matching score, the higher the matching degree of the hit image or the hit video segment with the first text. Optionally, the hit images and the target videos can be sorted in descending order of the matching scores.

[0009] The visual media search method provided in the first aspect, first of all, this method can not only be applied to image search, but also be applied to video search, providing users with more and more comprehensive search results and improving the user experience. Secondly, this method can sort and display the images and videos in the search results together, reflecting the matching degrees of different images or videos with the search text, which is more conducive to users to find the images or videos that meet the search intention and improves the user experience.

[0010] Combined with the first aspect, in some implementation manners of the first aspect, at least one hit image and at least one target video are sorted according to the matching scores of each hit result to obtain a first sorting result, including: for the first target video, the highest score among the matching scores of all the first hit video segments is used as the matching score of the first target video; the first target video is any one of the at least one target video, and the first hit video segment is the hit video segment in the first target video; at least one hit image and at least one target video are sorted according to the matching scores to obtain a first sorting result.

[0011] That is to say, for any target video, the highest matching score of each hit video segment included in the target video is used as the matching score of the target video to participate in the sorting.

[0012] A target video may include multiple hit video segments, and the corresponding matching scores of different hit video segments may be different. In this implementation manner, using the highest matching score of all the hit video segments as the matching score of the target video can better reflect the matching degree of the target video with the first text. The target video participates in the matching degree sorting based on this matching score, and the obtained first sorting result is more accurate, which can further improve the user experience.

[0013] In a possible implementation manner, the number of the first hit video segments is multiple, and the method further includes: sorting the multiple first hit video segments according to the matching scores to obtain a second sorting result.

[0014] As described above, a target video may include multiple hit video segments. Multiple hit video segments of the same target video may be sorted, and the sorting result is called the second sorting result.

[0015] In a possible implementation, after displaying at least one hit image and at least one target video according to the first sorting result, the method further includes: receiving a second operation of the user to play the first target video; in response to the second operation, sequentially playing multiple first hit video segments according to the second sorting result.

[0016] For example, a certain target video includes hit video segment 1, hit video segment 2, and hit video segment 3. The sorting result obtained by sorting these three hit video segments from high to low according to the matching score is: hit video segment 2, hit video segment 3, and hit video segment 1. Then, when the user clicks to play this target video, these three hit video segments can be played in this order. In this way, hit video segments with a high degree of matching with the search text are preferentially shown to the user, which is more likely to match the user's search intention and improves the user experience.

[0017] In a possible implementation, displaying at least one hit image and at least one target video includes: displaying thumbnails of each hit image and thumbnails of each target video; wherein, the thumbnail of the first target video is the thumbnail of a frame image in the second hit video segment, and the second hit video segment is the video segment with the highest matching score among all the first hit video segments; one or more of the following data are displayed in the thumbnail of the first target video: the start playing moment of the second hit video segment, the end playing moment of the second hit video segment, the middle playing moment of the second hit video segment, the total number of the first hit video segments, and the total number of video segments included in the first target video.

[0018] Optionally, the thumbnail of the first target video may be a frame image in the hit video segment with the highest matching score, for example, the representative frame of the hit video segment with the highest matching score. This representative frame is likely to match the search text (the first text), and the user can intuitively know that the target video contains the content he wants to search through the thumbnail, improving the user experience.

[0019] In addition, further displaying the above data in the thumbnail facilitates the user to intuitively know the information of the video and the information of the hit video segment, further improving the user experience.

[0020] In a possible implementation, the matching scores of each hit result are determined respectively, including: determining multiple matching dimensions according to the first text; determining the dimension scores of each hit result on each matching dimension, where the dimension score of the first hit result on the first matching dimension represents the matching degree between the first hit result and the first text on the first matching dimension; the first hit result is any one of at least one hit result, and the first matching dimension is any one of multiple matching dimensions; calculating the weights of each matching dimension according to the dimension scores of each hit result on each matching dimension; and performing weighted summation on the dimension scores of the first hit result on each matching dimension according to the weights of each matching dimension to obtain the matching score of the first hit result.

[0021] Optionally, the first text can be recognized for natural language understanding, and various information categories such as recognized semantic entities and person relationships are used as matching dimensions respectively. At the same time, the semantic vector is used as a matching dimension.

[0022] In this implementation, the matching score is calculated based on the weighted summation method, which fully considers the contribution of information in different matching dimensions to the final matching degree. Therefore, the obtained matching score is more in line with the actual situation, has higher accuracy, and the probability of the same matching scores for different hit results is small, reducing the occurrence of situations where the matching degree of hit results cannot be distinguished. In addition, this method does not require building a data set, so it does not need to collect user data, can better protect the privacy and security of users, meet the "minimization" principle of user data, and improve the user experience.

[0023] In a possible implementation, calculating the weights of each matching dimension according to the dimension scores of each hit result on each matching dimension includes: performing normalization processing on the dimension score matrix to obtain a normalized matrix, where the dimension score matrix is a matrix composed of the dimension scores of each hit result on each matching dimension; calculating the information entropy of each matching dimension based on the normalized matrix; and calculating the weights of each matching dimension according to the information entropy of each matching dimension, where the weight of the first matching dimension is negatively correlated with the information entropy of the first matching dimension.

[0024] In this implementation, the weights are determined based on the information entropy by calculating the information entropy of each matching dimension. There is no need to set the weights manually. Moreover, the information entropy can quantitatively measure the magnitude of the data variation degree of different matching dimensions and reflect the contribution degree of the data of different matching dimensions to the final matching score. Therefore, the weights of different matching dimensions can be accurately allocated to improve the accuracy of matching score calculation. In addition, this method performs normalization processing on the dimension score matrix and calculates the information entropy and weights based on the normalized matrix obtained from the normalization processing, which can ensure that the matching degrees of hit results on different matching dimensions have a unified dimension and improve the accuracy of weight and matching score calculation.

[0025] In a second aspect, the present application provides a visual media search method, which is executed by an electronic device. The method includes: displaying a first interface of a gallery application; the first interface includes a search box, the search box includes a first text, and the first interface further includes a first thumbnail of a first target video; a first playback moment is displayed in the first thumbnail, and the first playback moment corresponds to a first target frame image in the first target video; in response to a triggering operation by the user on the first thumbnail, playing the first target video starting from the first playback moment; displaying a playback progress bar; the playback progress bar displays first marker information marking the first playback moment and second marker information marking a second playback moment, and the second playback moment corresponds to a second target frame image in the first target video; wherein, both the first target frame image and the second target frame image match the first text.

[0026] The first text is the search text input by the user through the search box. The first target video is a video that matches the first text, also known as a hit video.

[0027] The triggering operation on the first thumbnail can be, for example, clicking on the first thumbnail. The first marker information and the second marker information can be video tags, or can be information such as text or symbols.

[0028] Simply put, the first target video contains at least two frame images that match the first text. One frame is called the first target frame image, and the other frame is called the second target frame image. The playback moment corresponding to the first target frame image is the first playback moment, and the playback moment corresponding to the second target frame image is the second playback moment. When the user clicks on the thumbnail of the first target video, the electronic device starts playing the first target video from the first target frame image, that is, starts playing from a frame image that matches the first text. In this way, the user can see the picture that meets the search intention fastest and most directly, without the user having to search one by one through the playback progress bar, improving the user experience.

[0029] In addition, the first marker information and the second marker information are displayed in the playback progress bar, facilitating the user to intuitively know the position of the content related to the search text in the video, improving the user experience.

[0030] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: in response to an operation by the user on the second marker information, playing the first target video starting from the second playback moment.

[0031] For example, when the user clicks on the second marker information, the electronic device jumps the playback progress to the second playback moment, that is, jumps to the second target frame image. The second target frame image matches the first text. Therefore, with this method, the user can view the picture that matches the search text in one step, without the user having to drag the playback progress bar to find the relevant content, facilitating user operation and further improving the user experience.

[0032] In one possible implementation, the first target video also includes at least one third target frame image, the third target frame image is different from the first target frame image and the second target frame image, and the third target frame image matches the first text; the first playback time is the earliest one among the first playback time, the second playback time and the playback times corresponding to each third target frame image; or, the first target frame image is the one with the highest degree of matching with the first text among the first target frame image, the second target frame image and at least one third target frame image.

[0033] That is to say, the first target video contains multiple target frame images that match the first text. In this case, in one implementation, the thumbnail of the first target video (i.e., the first thumbnail) can display a target frame image with the earliest playback time. When the user clicks to play, the video starts playing from the target frame image with the earliest playback time. In other words, the target frame images that match the first text are played in chronological order, so that the display of the screen conforms to the chronological order, which is convenient for users to recall the relevant scenes of the video and improves the user experience.

[0034] As another implementation method, a target frame image with the highest degree of matching with the first text may also be displayed in the thumbnail of the first target video. When the user clicks to play, the video starts playing from the target frame image with the highest degree of matching with the first text. In this way, the images that are more consistent with the user's search intent are displayed first, thereby improving the user experience.

[0035] In a possible implementation, the playback progress bar also displays sorting information of the first target frame image, the second target frame image, and at least one third target frame image, and the sorting information represents the sorting of the matching degree between the frame image and the first text.

[0036] Optionally, the sorting sequence number may be displayed in the playback progress bar, for example, the numbers "1", "2", and "3" may be displayed. In this way, the user can intuitively know the matching degree between each target frame image and the search text, thereby improving the user experience.

[0037] In a possible implementation, the sorting information is represented by one or more of color, text, graphics, and symbols.

[0038] In a third aspect, the present application provides a visual media search method, which is executed by an electronic device. The method includes: displaying a first interface of a gallery application; the first interface includes a search box, the search box includes a first text, and the first interface further includes a first thumbnail of a first target video; a first playback moment is displayed in the first thumbnail; in response to a trigger operation by the user on the first thumbnail, playing the first target video starting from the first playback moment; displaying a playback progress bar; the playback progress bar displays marking information of a first playback time period and marking information of a second playback time period, the first playback time period corresponds to a first video segment in the first target video, the first playback time period includes the first playback moment; the second playback time period corresponds to a second video segment in the first target video; wherein, there is at least one frame image in the first video segment that matches the first text, and there is at least one frame image in the second video segment that matches the first text.

[0039] Briefly speaking, the first target video contains at least two video segments (referred to as hit video segments) that match the first text. One of them is called the first video segment, and the other is called the second video segment. When the user clicks on the thumbnail of the first target video, the electronic device starts playing the first target video from a certain frame image in the first video segment, that is, starts playing from a video segment that matches the first text. In this way, the user can see the video segment that meets the search intention fastest and most directly, without the need for the user to search one by one through the playback progress bar, improving the user experience.

[0040] In combination with the third aspect, in some implementation manners of the third aspect, playing the first target video starting from the first playback moment includes: playing the first video segment starting from the first frame image of the first video segment; or, playing the first video segment starting from a first target frame image in the first video segment, and the first target frame image matches the first text. Optionally, the first target frame image can be, for example, a representative frame of the first video segment.

[0041] In other words, when playing the hit video segment, it can start playing from the first frame image of the hit video segment. The video segment can be the result of dividing the video according to the scene. Therefore, there may be multiple frame images that match the first text in the hit video segment. This method can more comprehensively display the picture that meets the search intention to the user, improving the user experience.

[0042] Optionally, it can also start playing from the target frame image that matches the first text in the hit video segment. In this way, it intuitively shows the picture that meets the search intention to the user, improving the user experience.

[0043] In a possible implementation manner, the method further includes: after the first video segment is played, playing the second video segment.

[0044] That is to say, it is possible to only play the hit video segments and skip the non-hit video segments, which can save the user's time and improve the user experience.

[0045] In a possible implementation, the playback progress bar also displays first marker information corresponding to the first video segment and second marker information corresponding to the second video segment; the method further includes: in response to an operation by the user on the second marker information, playing the second video segment.

[0046] For example, when the user clicks on the second marker information, the electronic device jumps the playback progress to the second video segment. Since there is at least one frame image in the second video segment that matches the first text, with this method, the user can view the video segment that matches the search text in one step, without the need for the user to drag the playback progress bar to find relevant content, which is convenient for the user to operate and further improves the user experience.

[0047] In a possible implementation, playing the second video segment includes: starting to play the second video segment from the first frame image of the second video segment; or starting to play the second video segment from a second target frame image in the second video segment, where the second target frame image matches the first text.

[0048] The beneficial effects of this implementation are similar to those of the implementation of playing the first video segment and will not be elaborated here.

[0049] In a possible implementation, the playback progress bar also displays sorting information of the first video segment and the second video segment, and the sorting information represents the sorting of the video segments according to the degree of matching with the first text.

[0050] In a possible implementation, the sorting information is represented by one or more of color, text, graphics, and symbols.

[0051] For the beneficial effects of the above two implementations, refer to the method in the third aspect and details will not be repeated here.

[0052] In a fourth aspect, the present application provides a device, which is included in an electronic device and has the function of implementing the behaviors of the electronic device in the above first aspect and the possible implementation manners of the first aspect. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions. For example, a receiving module or unit, a processing module or unit, etc.

[0053] In a fifth aspect, the present application provides a device, which is included in an electronic device and has the function of implementing the behaviors of the electronic device in the above second aspect and the possible implementation manners of the second aspect.

[0054] Sixth aspect, the present application provides a device, which is included in an electronic device and has the function of implementing the behavior of the electronic device in the above-mentioned third aspect and its possible implementation manners.

[0055] Seventh aspect, the present application provides an electronic device, which includes: a memory, a processor, and a display screen; the display screen is used to display the user interface of the photo gallery; the memory is coupled to the processor, and the memory is used to store computer program code, the computer program code includes computer instructions, and the processor calls the computer instructions to enable the electronic device to execute any one of the technical solutions in the first aspect, the second aspect, or the third aspect.

[0056] Eighth aspect, the present application provides a chip, which includes a processor. The processor is used to read and execute the computer program stored in the memory to execute the method in the first aspect and its any possible implementation manners, or execute the method in the second aspect and its any possible implementation manners, or execute the method in the third aspect and its any possible implementation manners.

[0057] Optionally, the chip further includes a memory, and the memory is connected to the processor through a circuit or a wire.

[0058] Further optionally, the chip further includes a communication interface.

[0059] Ninth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the processor is enabled to execute any one of the technical solutions in the first aspect, the second aspect, or the third aspect.

[0060] Tenth aspect, the present application provides a computer program product, which includes: computer program code. When the computer program code runs on an electronic device, the electronic device is enabled to execute any one of the technical solutions in the first aspect, the second aspect, or the third aspect. Description of the Drawings

[0061] Figure 1 is a schematic diagram of an application scenario of a visual media search method provided by an embodiment of the present application;

[0062] Figure 2 is a schematic diagram of another application scenario of a visual media search method provided by an embodiment of the present application;

[0063] Figure 3 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application;

[0064] Figure 4 is a software structure block diagram of an electronic device provided by an embodiment of the present application;

[0065] Figure 5 It is a schematic diagram of the interface change of an example visual media search method provided by an embodiment of the present application;

[0066] Figure 6 It is a schematic diagram of an example search result interface provided by an embodiment of the present application;

[0067] Figure 7 It is another schematic diagram of an example search result interface provided by an embodiment of the present application;

[0068] Figure 8 It is another schematic diagram of an example search result interface provided by an embodiment of the present application;

[0069] Figure 9 It is another schematic diagram of an example search result interface provided by an embodiment of the present application;

[0070] Figure 10 It is a schematic diagram of the interface during the video playback process provided by an embodiment of the present application;

[0071] Figure 11 It is another schematic diagram of the interface during the video playback process provided by an embodiment of the present application;

[0072] Figure 12 It is another schematic diagram of the interface during the video playback process provided by an embodiment of the present application;

[0073] Figure 13 It is a schematic diagram of an example search entry provided by an embodiment of the present application;

[0074] Figure 14 It is a schematic diagram of the principle of an example visual media search method provided by an embodiment of the present application;

[0075] Figure 15 It is a process interaction diagram of an example visual media search method provided by an embodiment of the present application;

[0076] Figure 16 It is another process interaction diagram of an example visual media search method provided by an embodiment of the present application;

[0077] Figure 17 It is a process schematic diagram of an example video segmentation process provided by an embodiment of the present application;

[0078] Figure 18 It is a schematic diagram of the principle of an example video segmentation process provided by an embodiment of the present application;

[0079] Figure 19 It is a process schematic diagram of an example multi-way recall process provided by an embodiment of the present application;

[0080] Figure 20It is a schematic flowchart of a weighted fusion sorting process provided by an embodiment of the present application;

[0081] Figure 21 It is a schematic flowchart of a process for determining a matching score provided by an embodiment of the present application. Detailed implementation manners

[0082] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; herein, "and / or" is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality" means two or more than two.

[0083] Hereinafter, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include one or more of such features.

[0084] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in the specification of the present application are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0085] To better understand the embodiments of the present application, the following explains the terms or concepts that may be involved in the embodiments.

[0086] 1. Visual media

[0087] Visual media refers to the media that transmits information through vision. In the embodiments of the present application, visual media may include images, videos, etc.

[0088] 2. Semantic entity

[0089] The semantic entity can also be referred to as an entity. Semantic entities include, but are not limited to, entities with specific meanings such as time, person names (PER), and place names (LOC). Optionally, semantic entities can be identified through named entity recognition (NER) technology.

[0090] 3. Text semantic vector

[0091] A text semantic vector refers to a vector that can represent the semantics of an entire text.

[0092] 4. Visual semantic vector

[0093] A visual semantic vector refers to a vector that can represent the visual semantics of a visual medium. In the embodiments of this application, a visual semantic vector refers to a vector that can represent the semantics of an image or the semantics of a video segment.

[0094] 5. Vector similarity

[0095] Vector similarity is used to describe the similarity between two vectors. Generally, vector similarity can be calculated through the cosine similarity calculation formula. Of course, it can also be calculated through other methods.

[0096] 6. Visual content related and visual content unrelated

[0097] Visual content refers to the targets presented by a visual medium and the relationships between them. Computer vision enables a computer to have capabilities similar to human vision, including perceiving, understanding, analyzing, and interpreting visual content. Currently, the Generative Pre-trained Transformer 4 (GPT-4) can support inputting an image into the model and then outputting natural human language describing the important information in the image.

[0098] In the context of visual media search in this solution, in a set of texts, the information used to describe the visual content in a visual medium is called "visual content related" information. Visual content related information can only be obtained by an electronic device through a text semantic understanding model. In a set of texts, the information unrelated to the visual content of a visual medium is called "visual content unrelated" information. Visual content unrelated information can be obtained through user annotation, pre-recorded by an electronic device, and identified through preset algorithm analysis. In the embodiments of this application, visual content unrelated information can include locations, times, names, file attributes, person names, tags, person relationship information, etc. Among the above-listed visual content unrelated information, obtaining the time, name, and file attributes can be automatically recorded by the electronic device when obtaining the visual medium. Person names and tags can be annotated by users. Person relationship information can be identified through person relationship analysis algorithms.

[0099] For example, in the sentence "The sky photographed in Beijing this year", the information such as "this year" (time) and "Beijing" (location) is not related to the visual content of the visual media. Therefore, "this year" and "Beijing" are information unrelated to the visual content. The object corresponding to "sky" can be perceived by the human visual ability and is information used to describe the visual content in the visual media. Thus, it is information related to visual semantics.

[0100] Before elaborating on the visual media search method provided by the embodiments of the present application in detail, first, the application scenario of this method and the structure of the applicable electronic device will be described.

[0101] Optionally, the visual media search method provided by the embodiments of the present application can be applied to the gallery application (APP) of an electronic device to search for images and videos that match the search content input by the user. The following will be described in conjunction with the interface diagram.

[0102] Exemplarily, Figure 1 FIG. is a schematic diagram of an application scenario of a visual media search method provided by an embodiment of the present application. Taking the electronic device as a mobile phone as an example, as Figure 1 shown in FIG. (a), the main interface of the mobile phone includes the icon 101 of the gallery APP. The user clicks on this icon 101. In response to the user's click, the mobile phone opens the gallery APP and displays Figure 1 the album interface shown in FIG. (b). The album interface 102 includes a search box 103. The user can enter the search content in the search box 103. For example, the user can enter "Beijing" in the search box. In response to the user's operation, the mobile phone can search for images with the shooting location of "Beijing" in the gallery and display the search results, as Figure 1 shown in FIG. (c).

[0103] In the related art, the search function is only applicable to the search for images and not applicable to the search for videos. Specifically, as Figure 1 shown in FIG. (c), the search results do not include videos.

[0104] Moreover, in the related art, the search is mainly based on the shooting location, shooting time, and attribute tags marked by the user for the images. Content without marked attribute tags cannot be searched, so the accuracy of the search results is not high and the user experience is poor. For example, as Figure 2 shown in FIG. (a), when the user enters the complex search text "a boy standing outdoors" in the search box 103 of the album interface 102, although there are images or videos in the gallery that match this search statement (as Figure 1 shown in FIG. (c)), however, the mobile phone cannot search for the results, as Figure 2as shown in Figure (b) therein. The user still needs to manually operate the scroll bar of the sliding picture gallery to search, which is cumbersome and thus affects the user experience.

[0105] In addition, in the related art, when searching for images based on shooting time, shooting moment, etc., the search results do not involve sorting problems. If the user needs to find images that meet the actual search intention from these search results, the user needs to open the search results and search one by one, which is cumbersome and brings a poor search experience to the user. For example, when the user wants to search for images related to "boys standing outdoors", through the Figure 2 process shown in the figure cannot be searched, so it can only be searched by shooting time or shooting moment. Assuming that the images related to "boys standing outdoors" were taken in August 2022, the user can continue to enter "August 2022" in the search box 103 in Figure 2 Figure (b) therein. In response to the user's operation, the mobile phone displays the search results, as shown in Figure 2 Figure (c) therein. The search results include 58 images. The user needs to click on the "More" option 201 to enter the interface shown in Figure 2 Figure (d) therein and find the images related to "boys standing outdoors" from it.

[0106] In view of this, the embodiments of the present application provide a visual media search method. On the one hand, this method can not only be applied to images, but also to videos, providing more and more comprehensive search results to improve the user experience. On the other hand, this method is based on semantic vectors for search and matching. Therefore, the user can perform fuzzy search of sentences. For example, the user can input search sentences such as "boys standing outdoors", "women doing yoga", etc., and the electronic device can search based on this sentence, and can more accurately search for visual media that meets the actual intention of the user, improving the search fineness and the user experience. On the third hand, this method can sort and display the images and videos in the search results together, reflecting the matching degree of different images or videos with the search sentence, which is more conducive to the user to find images or videos that meet the search intention and improves the user experience.

[0107] The visual media search method provided by the embodiments of this application can be applied to electronic devices such as mobile phones, tablet computers, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs) that can install application programs (APPs). The embodiments of this application do not impose any restrictions on the specific types of electronic devices.

[0108] Exemplarily, Figure 3 is a schematic structural diagram of an electronic device 100 provided by the embodiments of this application. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0109] It can be understood that the structure schematically shown in the embodiments of this application does not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0110] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0111] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.

[0112] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0113] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0114] It can be understood that the interface connection relationships among the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0115] The charging management module 140 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 can receive the charging input from a wired charger through the USB interface 130. In some embodiments of wireless charging, the charging management module 140 can receive the wireless charging input through the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device through the power management module 141.

[0116] The electronic device 100 can implement a shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, the application processor, etc.

[0117] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present application, taking the Android system with a layered architecture as an example, the software structure of the electronic device 100 is illustratively described.

[0118] Figure 4 It is the software structure block diagram of the electronic device 100 in the embodiments of the present application. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely the application layer, the application framework layer, the Android runtime and the system libraries, and the kernel layer.

[0119] The application layer may include a series of application packages. As Figure 3 shown, the application packages may include a gallery APP, a search module, a natural language understanding module, and computer vision (CV) services, etc.

[0120] The gallery APP is also called the gallery service or the gallery service module, etc., and is used to provide services such as storage, management, display, and recommendation of visual media such as images and videos to users. Optionally, the gallery APP may have a database, hereinafter referred to as the gallery database. The gallery database is used to store relevant data of the gallery APP, such as images and their attribute information, videos and their attribute information, visual semantic vectors, person relationship information, and so on.

[0121] Optionally, the gallery APP may have multiple functions, such as an album function, a one - click blockbuster function, etc. The data in the gallery database can be applied to each function of the gallery APP, and the data returned by the search module to the gallery APP can also be applied to any function of the gallery APP. This application does not make any limitations in this regard.

[0122] The search module is used to search for visual media according to the content input by the user. Optionally, the search module may have its own database for storing index data. Among them, the indexing method of the index data can be an inverted index. Hereinafter, the database of the search module will be referred to as an inverted index library.

[0123] The natural language understanding module is used to understand the information contained in the text. In the embodiments of this application, the natural language understanding module can be used to identify semantic entities, person relationships, category information, etc. in the text. Category information refers to information representing the type to which an image or video belongs. Exemplarily, the category information can be a person, a plant, an animal, a building, or a natural scenery, etc.

[0124] In one embodiment, as Figure 3 shown, the CV service may include a person relationship analysis module, a video segmentation module, and a multi - modal understanding module, etc.

[0125] The person relationship analysis module is used to analyze the social relationships between the people contained in the visual media to obtain person relationship information.

[0126] The video segmentation module is used to perform segmentation processing on the video, divide the video into multiple video segments according to the semantic content, and extract one frame of image from each video segment as a representative frame (also referred to as a representative frame image). The representative frame is used to represent the semantics of the video segment.

[0127] The multi - modal understanding module, also known as the multi - modal semantic understanding module, the multi - modal semantic understanding model, etc., is used to perform semantic understanding on one or more of various modal information such as text, image, video, audio, etc. In the embodiments of this application, the multi - modal understanding module can be used to perform semantic understanding on an image to obtain a visual semantic vector, and can also be used to perform semantic understanding on text (such as search text) to obtain a text semantic vector.

[0128] Of course, in addition to Figure 4 the various modules or APPs shown, the application layer may also include applications such as a camera, a call, a map, a Bluetooth, a short message, etc. (not shown in the figure). This application does not make any limitations in this regard.

[0129] The application framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer includes some predefined functions.

[0130] The application framework layer may also include a window manager, a content provider, a view system, a telephone manager, a resource manager, a notification manager, etc.

[0131] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.

[0132] The content provider is used to store and obtain data, and make this data accessible to application programs. The data may include videos, images, audio, dialed and answered calls, browsing history and bookmarks, phone books, etc.

[0133] The view system includes visible controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build application programs. The display interface can be composed of one or more views. For example, a display interface including a short message notification icon may include a view for displaying text and a view for displaying pictures.

[0134] The telephone manager is used to provide the communication function of the electronic device 100. For example, the management of call states (including connection, disconnection, etc.).

[0135] The resource manager provides various resources for application programs, such as localized strings, icons, pictures, layout files, video files, etc.

[0136] The notification manager enables application programs to display notification information in the status bar, can be used to convey notification-type messages, and can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that the download is completed, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or a scroll bar text, such as the notification of a background-running application program, and can also be a notification that appears on the screen in the form of a dialogue window. For example, prompt text information in the status bar, emit a prompt sound, the electronic device vibrates, the indicator light flashes, etc.

[0137] Optionally, each application package in the application layer can implement the related functions of the gallery APP by calling the algorithms or modules of the application framework layer. For example, when the face analysis module performs face clustering, it can call the face clustering algorithm module (not shown in the figure) of the application framework layer to cluster the faces in the image to obtain a face clustering result. Another example is that when the person relationship analysis module performs person relationship analysis, it can call the social circle recognition algorithm module (not shown in the figure) of the application framework layer, and based on the face clustering result, identify the central person and discover the social circle of the central person. The detailed implementation of the above process will be further elaborated in the subsequent embodiments.

[0138] The Android runtime includes the core libraries and the virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0139] The core libraries consist of two parts: one part is the functional functions that the Java language needs to call, and the other part is the core libraries of Android.

[0140] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files in the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.

[0141] The system libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (such as OpenGL ES), 2D graphics engines (such as SGL), etc.

[0142] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.

[0143] The media libraries support the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media libraries can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0144] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.

[0145] The 2D graphics engine is the graphics engine for 2D drawing.

[0146] The kernel layer is the layer between the hardware and the software. The kernel layer includes at least a display driver, a camera driver, an audio driver, and a sensor driver.

[0147] For the sake of easy understanding, the following embodiments of this application will take an electronic device with Figure 3 and Figure 4 the structure shown as an example, and in combination with the accompanying drawings and application scenarios, specifically elaborate on the visual media search method provided by the embodiments of this application.

[0148] First, the interface effect of the visual media search method provided by the embodiments of this application will be described.

[0149] Exemplarily, Figure 5 is a schematic diagram of the interface change of an example of the visual media search method provided by the embodiments of this application. As Figure 5As shown in Figure (a), the album interface 102 of the gallery APP is displayed on the mobile phone screen. The album interface 102 includes a search box 103, and the user enters the search text "boys standing outdoors" in the search box 103. In response to the user's operation, the mobile phone searches the gallery database for images and videos that match "boys standing outdoors" and displays the search results, as shown in Figure 5 the search result interface 501 shown in Figure (b). The search result interface 501 may include a first display area 502, a second display area 503, and a third display area 504. In the first display area 502, some of the search results (e.g., 8) are shown, including images and videos. In the second display area 503, a video that best matches the search text in the search results is shown, and the total number of videos included in the search results, "4", is shown. In the third display area 504, an image that best matches the search text in the search results is shown, and the total number of images included in the search results, "6", is shown. It can be understood that in the embodiments of the present application, the search results can be shown in the form of thumbnails, that is, the thumbnails of the images are displayed, or the thumbnail of a certain frame image of the video is used as the thumbnail of the video for display.

[0150] It can be seen that in the embodiments of the present application, when the search content is a sentence, visual media that semantically matches the sentence can also be obtained, that is, the method can perform fuzzy search based on the sentence. Moreover, in the search results, in addition to images, videos are also included.

[0151] Specifically, for videos, the method pre-segments each video, and divides the video into multiple video segments according to semantic content. During the search, the search text is matched with the semantics of each video segment. If there is at least one video segment in a certain video that semantically matches the search text, then that video is shown as a search result. In subsequent embodiments, the video segment in the video that matches the search text is called a hit video segment.

[0152] In one embodiment, when showing the search results, the videos and images in the search results can also be sorted according to the degree of match with the search text. Optionally, the sorting principle can be: the higher the degree of match with the search text, the higher the display order. Among them, the degree of match between an image or a video and the search text can be characterized by a match score, and the higher the match score, the higher the degree of match. For example, Figure 5 in Figure (b), the search results from left to right and top to bottom include video 5051, image 5052, video 5053, image 5054, image 5055, image 5056, image 5057, and image 5058 in turn, and the match scores of these videos or images decrease in turn.

[0153] Among them, the matching degree between the video and the search text can be characterized by the hit video segment with the highest matching score in the video. For example, Figure 5 in Figure (b) of Figure 5 , the video 5051 includes three hit video segments. Among them, the first hit video segment has the highest matching score, so this matching score is used as the score of the video 5051 to compare and sort with the scores of other videos or images.

[0154] In addition, it can be understood that when the user clicks the "More" control 506, the electronic device responds to the user's operation and displays all search results. All search results can also be sorted and displayed according to the matching degree with the search text, which will not be elaborated here.

[0155] The display form of the videos in the search results will be described below.

[0156] In one embodiment, the cover of the video can display relevant information of the video. For example, as Figure 5 shown in Figure (b) of Figure 5 , the cover of the video 5051 can display the total duration of the video, which is 8 minutes and 32 seconds (08:32). The cover of the video 5053 can display the total duration of the video, which is 14 minutes and 18 seconds (14:18).

[0157] In one embodiment, the cover of the video in the search results can also display the start playing time (abbreviated as start time), end playing time (abbreviated as end time), middle playing time (abbreviated as middle time), etc. of a certain hit video segment. For example, when the video includes three hit video segments, the cover of the video can display the start time of the first hit video segment.

[0158] In another embodiment, the start time, end time or middle time of the hit video segment with the highest matching score can also be displayed on the video cover. As Figure 5 shown in Figure (b) of Figure 5 , taking the start time of the highest-score hit video segment in the video 5051 as 2 minutes and 18 seconds (02:18) and the end time as 3 minutes and 18 seconds (03:18) as an example, the cover of the video 5051 can display the start time 2 minutes and 18 seconds (02:18), or the end time 3 minutes and 18 seconds (03:18), or the middle time 2 minutes and 48 seconds (02:48).

[0159] In yet another embodiment, after the video is divided into video segments, a frame image can be selected from each video segment as a representative frame, and the playing time corresponding to the representative frame of the hit video segment can also be displayed on the cover of the video. For example, if the playing time corresponding to the representative frame of the hit video segment is 2 minutes and 0 seconds (02:00), the cover of the video can display this time.

[0160] In one embodiment, the cover of a video can be a thumbnail of any frame image in the video. For example, the first frame image of the video, or the first frame image of the hit video segment with the highest matching score, or the middle frame of the hit video segment with the highest matching score, or the representative frame of the hit video segment with the highest matching score, etc. For example, as Figure 5 shown in figure (b) of

[0161] The above embodiments are all described by taking the display of information of one hit video segment on the video cover in the search results as an example. In other embodiments, the video cover in the search results can also display information of multiple hit video segments.

[0162] Exemplarily, Figure 6 is a schematic diagram of a search result interface provided by an embodiment of the present application. As Figure 6 shown, the video cover can display the total number of hit video segments and the total number of video segments included in the video, so that the user can clearly know the number of hit video segments in the video. For example, video 5051 is divided into 5 video segments in total, and 3 of them match the search statement "boys standing outdoors", that is, it includes 3 hit video segments. Then, the cover of video 5051 can display information such as "3 hits / 5 segments in total". The same is true for video 5053 and will not be elaborated here.

[0163] In another embodiment, the video cover can also display one or more of the start time, end time, middle time, or the playing time corresponding to the representative frame of multiple hit video segments, so that the user can clearly know the positions of the hit video segments in the video. For example, the start times of multiple hit video segments can be displayed, or the playing times corresponding to the representative frames of multiple hit video segments can be displayed.

[0164] Suppose the three hit video segments of video 5051 are the first video segment, the third video segment, and the fifth video segment respectively. Among them, the start time of the first video segment is 2 minutes and 0 seconds (02:18), the end time is 3 minutes and 18 seconds (03:18), and the playing time corresponding to the representative frame is 2 minutes and 10 seconds (02:10); the start time of the second video segment is 5 minutes and 22 seconds (05:22), the end time is 6 minutes and 40 seconds (06:40), and the playing time corresponding to the representative frame is 5 minutes and 45 seconds (05:40); the start time of the third video segment is 8 minutes and 0 seconds (08:00), the end time is 8 minutes and 32 seconds (08:32), and the playing time corresponding to the representative frame is 8 minutes and 12 seconds (08:12).

[0165] Exemplarily, Figure 7 is another schematic diagram of a search result interface provided by an embodiment of the present application. Optionally, refer to Figure 7, the cover of video 5051 can display information such as "Hit 02:18 - 03:18", "Hit 05:22 - 06:40", and "Hit 08:00 - 08:32". The same applies to video 5053 and will not be elaborated here.

[0166] Exemplarily, Figure 8 is another schematic diagram of the search result interface provided by an embodiment of the present application. Optionally, refer to Figure 8 , the cover of video 5051 can display information such as "Hit 02:10", "Hit 05:40", and "Hit 08:12". The same applies to video 5053 and will not be elaborated here.

[0167] It can be understood that when the number of hit video segments is large, the video cover can display information of N video segments with relatively high matching scores among them. N can be, for example, 3, or 2, etc. In this way, it can prevent too much information from blocking the video cover and affecting the user experience.

[0168] In one embodiment, the cover of the video can also display frame images of multiple hit video segments. For example, representative frames of N video segments with relatively high matching scores can be spliced into an image as the cover of the video, so that users can intuitively see the content in the hit video segments and improve the user experience. Exemplarily, Figure 9 is another schematic diagram of the search result interface provided by an embodiment of the present application. Refer to Figure 9 , the representative frames of the first hit video segment, the second hit video segment, and the third hit video segment of video 5051 can be spliced and used as the cover of video 5051. The same applies to video 5053 and will not be elaborated here.

[0169] In summary, the method provided by the embodiments of the present application can have multiple display methods for search results, and the embodiments of the present application do not make any limitations in this regard.

[0170] The following describes the playing process of the video in the search results.

[0171] Exemplarily, Figure 10 is a schematic diagram of the interface of a video playing process provided by an embodiment of the present application. As Figure 10 shown in figure (a) of, when the user clicks on video 5051 in the search results, in response to the user's operation, a video playing interface 1001 is displayed. The video playing interface 1001 displays video 5051 and a playing control 1002. The cover of video 5051 can have the same display effect as the cover described in the above embodiments.

[0172] The user clicks on the play control 1002. In response to the user's operation, the mobile phone starts playing the video 5051. Optionally, the mobile phone can play each hit video segment in the order of the playback time period, that is, first play the first hit video segment (02:18 - 03:18), then play the second hit video segment (05:22 - 06:40), and then play the third hit video segment (08:00 - 08:32). After the three hit video segments are played, the playback stops. Taking the image displayed on the video cover as the first frame image of the first hit video segment as an example, after the user clicks on the play control 1002, the playback interface of the video 5051 can be as Figure 10 shown in Figure (b) of. It can be seen that after the video starts playing, a playback progress bar 1003 can be displayed on the interface. Currently, the video starts playing from the first hit video segment, so the playback time displayed in the playback progress bar 1003 is 02:18.

[0173] Optionally, the mobile phone can also play in the order of the matching scores of the hit video segments from high to low. For example, assuming that the matching score of the second hit video segment > the score of the first hit video segment > the score of the third hit video segment, then play in the order of the second hit video segment, the first hit video segment, and the third hit video segment. After the three hit video segments are played, the playback stops.

[0174] In addition, it can be understood that when playing each hit video segment, it can start playing from the first frame of the hit video segment, or start playing from the middle frame of the hit video segment, or start playing from the representative frame of the hit video segment. The representative frame can best represent the content of the video segment. Therefore, starting to play the hit video segment from the representative frame enables the user to see the picture related to the search text in the video segment fastest and most directly, improving the user experience.

[0175] Exemplarily, Figure 11 is a schematic diagram of the interface of another example of the video playback process provided by the embodiment of the present application. In some embodiments, when the video is played, information about each hit video segment can be displayed in the playback progress bar 1003. For example, the position of each hit video segment in the video can be displayed in the playback progress bar 1003, as Figure 11 shown in Figure (a) of, where the hit video segments are shown as black long bars. Optionally, corresponding text can also be displayed above the playback progress bar 1003 to indicate the hit video segments. For example, the numbers "1", "2", "3" are displayed to indicate the three hit video segments; or according to the matching scores, "Match Rank 1", "Match Rank 2", "Match Rank 3", etc. are displayed to show the matching rank of the video segment with the search text.

[0176] In another embodiment, the matching scores of the respective hit video segments can also be distinguished by different colors. For example, the hit video segment with the highest matching score is marked in red in the progress bar, the hit video segment with the second highest matching score is marked in pink in the progress bar, and the hit video segment with a relatively low matching score is marked in green in the progress bar, as Figure 11 shown in Figure (b) of Figure 11 . It should be noted that in Figure (b) of

[0177] , different colors are represented by different filling patterns, which is only an example and does not represent the actual display effect. Figure 12 Exemplarily, Figure 12 Figure (a) of Figure 12 shows another schematic diagram of the interface during the video playback process provided by the embodiment of the present application. Optionally, marking information, such as a video label, can also be set at the hit video segment. As Figure 12 shown in Figure (a) of

[0178] , the video label is shown as a black inverted triangle. When the user clicks on the label, the video directly jumps to the corresponding hit video segment to start playing. As

[0179] shown in Figure (a) of Figure 13 , when the video is played to 02:25, if the user clicks on the video label of the second hit video segment (05:22 - 06:40) with the highest matching score, the hit video segment will start playing, as Figure 13 shown in Figure (b) of Figure 13 .

[0180] The following describes the implementation process of the visual media search method provided in the embodiments of the present application in combination with the interaction diagram corresponding to the above interface changes.

[0181] The method may include two stages: index construction and search. Among them, the index construction stage belongs to the offline stage, and the search stage belongs to the online stage.

[0182] Exemplarily, Figure 14 is a schematic diagram of the principle of a visual media search method provided in the embodiments of the present application. As Figure 14 shown, in the index construction stage, for the videos in the gallery database, after frame splitting processing, multiple frame images are obtained; the multiple frame images are segmented to obtain multiple video segments; representative frames are extracted from the video segments, and the content of the video segments is represented by the representative frames; for the images in the gallery database, no frame splitting and segmentation processing are required. Therefore, the images in the gallery database and the representative frames obtained after segmenting the videos are subjected to analysis of the relationship between characters to obtain character relationship information, and semantic understanding is performed on these images and representative frames (for example, inputting an image semantic understanding model) to obtain corresponding visual semantic vectors; an inverted index is constructed based on the obtained visual semantic vectors and character relationship information, and the inverted index is stored in the inverted index library.

[0183] In the search stage, natural language understanding is performed on the search text input by the user to identify semantic subjects, character relationships, category information, etc. contained in the search text; in addition, semantic understanding is performed on the search text (for example, inputting a text semantic understanding model) to obtain a corresponding text semantic vector; multi-way recall is performed from the inverted index library based on the semantic subject, character relationship, category information, and text semantic vector; then, the results of multi-way recall are filtered and sorted to obtain search results; the search results are displayed on the search interface.

[0184] Optionally, when the above image semantic understanding model and text semantic understanding model are trained, they can be trained in the direction of contrast learning to make the semantic understanding of the two more consistent and improve the accuracy of visual media search.

[0185] The following separately describes the specific implementation processes of the two stages.

[0186] 1. Index construction stage

[0187] Figure 15 is a process interaction diagram of a visual media search method provided in the embodiments of the present application. As Figure 15 shown, the index construction stage of the method may include:

[0188] S101. In response to the user adding and / or modifying visual media, the gallery APP stores the incremental visual media added and / or modified by the user and their attribute information in the gallery database.

[0189] Optionally, the user can add visual media by means such as shooting, downloading, taking screenshots, etc. In addition, the user can also modify the existing visual media to obtain new visual media. Such modifications include, but are not limited to: beautification, editing, custom addition of personal names, adding watermarks, etc.

[0190] The attribute information of visual media can include, but is not limited to: identification information, acquisition location, acquisition time, personal name, label (also known as classification label). Optionally, the identification information of visual media can be represented by an identity document (ID) (such as a hash ID), or can be represented by a name, path, etc. This application does not make any limitation in this regard. In the following embodiments, taking the identification information of visual media being represented by ID as an example for illustration.

[0191] The acquisition location refers to the location where the visual media is acquired. The acquisition time refers to the time when the visual media is acquired. For a captured video or image, the acquisition location can be the shooting location, and the acquisition time can be the shooting time. For a screenshot, the acquisition location can be the screenshot location, and the acquisition time can be the screenshot time. For a downloaded video or image, the acquisition location can be the download location, and the acquisition time can be the download time.

[0192] The personal name refers to the name added by the user to the person in the image or video. The label refers to the information characterizing the type or attributes of the image or video. Exemplarily, the label can be person, plant, animal, building, or natural scenery, etc.

[0193] For the target visual media newly added by the user and its attribute information, the gallery APP stores the newly added visual media and its attribute information in the gallery database. For the visual media and its attribute information modified by the user, the gallery APP synchronously modifies the visual media and its attribute information stored in the gallery database.

[0194] In some other embodiments, the gallery APP can store the newly added visual media and its attribute information in the cloud to relieve the storage pressure on the local of the electronic device.

[0195] For the sake of convenience of description, in the following embodiments, the visual media added and / or modified by the user is referred to as incremental visual media or target visual media, the video in the visual media added and / or modified by the user is referred to as incremental video or target video, and the image in the visual media added and / or modified by the user is referred to as incremental image or target image.

[0196] S102. When the electronic device is in a charging state and the screen is in a screen-off state, the gallery APP sends a segmentation request to the video segmentation module of the CV service.

[0197] The segmented request is used to request the segmented processing of the incremental video and obtain the representative frames of the video segments. Optionally, the incremental video or the ID of the incremental video may be carried in the segmented request.

[0198] Optionally, the gallery APP may subscribe to the charger plugging and unplugging status from the relevant modules in the electronic device, and when the charger plugging and unplugging status changes, the module sends a notification to the gallery APP. In this way, the gallery APP can know in real time whether the charger is in the plugged state, that is, know whether the electronic device is in the charging state. Similarly, the gallery APP may also subscribe to the screen status from the relevant modules in the electronic device, and when the screen status changes, the module sends a notification to the gallery APP. In this way, the gallery APP can know in real time whether the screen is in the off state.

[0199] In another embodiment, the gallery APP may also send a segmented request to the video segmentation module when the electronic device is in the charging state, the screen is in the off state, and the current time is within a preset time period. Optionally, the preset time period may be, for example, from 00:00 to 07:00 every day. In this way, it is possible to reduce the disturbance to the user and improve the user experience.

[0200] S103. In response to the segmented request, the video segmentation module performs segmented processing on the incremental video to obtain multiple video segments and their attribute information, and extracts representative frames from each video segment.

[0201] Specifically, the video segmentation module may call a preset video segmentation algorithm to perform segmented processing on each incremental video. The video segmentation by the video segmentation algorithm is mainly divided into three processes: the first process is to perform frame splitting on the incremental video to split the incremental video into multiple frame images; the second process is to segment the multiple frame images obtained by frame splitting based on the semantics of the video, and each video segment obtained by segmentation may respectively represent a scene, and there is a certain correlation between the frame images within each video segment; the third process is to extract one frame image from each video segment as the representative frame.

[0202] The representative frame is used to represent the semantics of the video segment. In one embodiment, the first frame image, the last frame image, or the middle frame image of the video segment may be used as the representative frame. In another embodiment, the representative frame may also be determined by calculating parameters such as the jitter degree, clarity, label, and pixel transformation degree of the frame images in the video segment. The embodiments of the present application do not make specific limitations on the method for determining the representative frame.

[0203] Optionally, the attribute information of the video segment may include video segment identification information, start time, end time, label union, etc. The identification information of the video segment may be, for example, the video segment ID. The label union refers to the union of the labels (referred to as frame labels) of the frame images in the video segment.

[0204] The specific implementation of step S103 will be further described in the subsequent embodiments.

[0205] S104. The video segmentation module returns the video segment, its attribute information, and the representative frame to the gallery APP.

[0206] S105. The gallery APP stores the video segment, its attribute information, and the representative frame in the gallery database.

[0207] S106. The gallery APP sends a person relationship analysis request to the person relationship analysis module in the CV service. The person relationship analysis request may carry the incremental image and the representative frame.

[0208] The person relationship analysis request is used to request the analysis of the person relationship information of the incremental image and the representative frame. The person relationship information is used to characterize the relationship type between the persons included in the image and the central person. The person relationship can be, for example, family member, son, daughter, parent, colleague, friend, etc. The central person is the person in the center position of the social relationship. The central person is generally considered to be the device owner or a person with a relatively high intimacy with the device owner.

[0209] S107. The person relationship analysis module analyzes the person relationship information included in the incremental image and the representative frame in response to the person relationship analysis request.

[0210] Specifically, the person relationship analysis module may call a relevant module to identify the incremental image, identify the human faces in the incremental image or the representative frame, and analyze the human faces, such as analyzing the age corresponding to the human face, the gender corresponding to the human face, and clustering the human faces included in multiple images to obtain a human face clustering result. Human face clustering is to cluster human faces with similar features into one class. In this way, at least one class is obtained, and each class may correspond to a person.

[0211] After that, the person relationship analysis module performs person relationship analysis based on the human face clustering result to obtain the person relationship information of the incremental image and the representative frame. Person relationship analysis is also called social relationship analysis, person relationship discovery, etc.

[0212] Optionally, on the one hand, the person relationship analysis module may obtain the person relationship information actively added by the user to the person. On the other hand, the person relationship analysis module may call a social circle recognition algorithm, based on the human face clustering result, identify the central person among the persons included in multiple images, and identify the social circle formed by the persons in multiple images and the type of the social circle based on the central person. The social circle of a person refers to the set of persons with social relationships. The type of the social circle is used to characterize the relationship between the persons in the social circle and the central person. Optionally, the type of the social circle may include family member, son, daughter, parent, colleague, friend, etc. That is to say, the type of the social circle may be used as the person relationship information.

[0213] Optionally, when identifying social circles, first, based on the face clustering results, a person-image heterogeneous graph can be generated, which is used to represent the correspondence between persons and images containing the persons. Then, based on the person-image heterogeneous graph, a co-occurrence image set is determined. Co-occurrence images refer to images with person co-occurrence relationships. Next, based on the attribute information of each image in the co-occurrence image set, the intimacy between pairwise persons in the face clustering results is determined. Based on the intimacy between pairwise persons, the person-image heterogeneous graph is simplified to a person connection graph. Based on the person-image heterogeneous graph, the central person is identified. Based on the person connection graph, community detection is performed to obtain at least one social circle.

[0214] When identifying the type of social circle, the person-image heterogeneous graph can be input into a graph neural network model and a multilayer perceptrons (MLP) model successively to obtain the relationship types between each person node and the central person node in the person-image heterogeneous graph. Then, based on the relationship types between each person node and the central person, the type of social circle is determined.

[0215] S108. The person relationship analysis module returns the incremental image and the person relationship information of the representative frame to the gallery APP.

[0216] S109. The gallery APP stores the incremental image and the person relationship information of the representative frame in the gallery database.

[0217] S110. The gallery APP sends an image semantic understanding request to the multi-modal understanding module in the CV service. The image semantic understanding request may carry the incremental image and the representative frame.

[0218] The image semantic understanding request is used to request semantic understanding of the incremental image and the representative frame.

[0219] S111. In response to the image semantic understanding request, the multi-modal understanding module performs semantic understanding on the incremental image and the representative frame to obtain the corresponding visual semantic vectors.

[0220] Specifically, the multi-modal understanding module may include a pre-trained image semantic understanding model. By inputting each incremental image and representative frame into the image semantic understanding model respectively, the corresponding semantic vectors can be obtained. In this embodiment, the semantic vectors corresponding to the images are referred to as visual semantic vectors.

[0221] S112. The multi-modal understanding module returns the visual semantic vectors of the incremental image and the representative frame to the gallery APP.

[0222] S113. The gallery APP stores the visual semantic vectors of the incremental image and the representative frame in the gallery database.

[0223] It can be understood that each visual semantic vector representing a frame is used to represent the semantics of the video segment to which it belongs, and has a unique correspondence with the video segment. Therefore, the visual semantic vector of the representative frame is also the visual semantic vector of the video segment.

[0224] S114. The gallery APP sends the attribute information, person relationship information, and visual semantic vectors of each video segment in the incremental images and incremental videos to the search module.

[0225] S115. The search module constructs an inverted index corresponding to the incremental visual media based on the attribute information, person relationship information, and visual semantic vectors of each video segment in the incremental images and incremental videos, and stores the inverted index in the inverted index library.

[0226] It can be understood that when constructing the index, the search module can also construct a forward index. In this embodiment, constructing an inverted index can improve the subsequent search efficiency.

[0227] Specifically, the attribute information, person relationship information, visual semantic vectors, etc. are collectively referred to as information. For incremental images, the search module can directly establish the correspondence between the incremental images and the information of the incremental images.

[0228] For incremental videos, the search module can establish the correspondence between the video segments and the information of the video segments, and establish the correspondence between each video segment and the incremental video to which it belongs and the attribute information of the incremental video to which it belongs. For the person relationship information and visual semantic vectors in the information of the video segment, the information of the representative frame can be used to represent them. That is to say, the person relationship information of the representative frame is used as the person relationship information of the video segment. The visual semantic vector of the representative frame is used as the visual semantic vector of the video segment.

[0229] In one embodiment, the relationship between the incremental video and its information is shown in Table 1 below. Optionally, the attribute information of the incremental video can be stored in a document, which can be called the parent document. The information of each video segment included in the incremental video can be stored in a document, which can be called the sub - document. An association relationship between the parent document and the sub - document is established.

[0230] Table 1

[0231]

[0232]

[0233] It should be noted that Table 1 is only an example for understanding the solution, and does not represent actual data, nor the form of actual data.

[0234] It can be understood that every time a user adds or modifies visual media, the process shown in the above steps S101 to S121 is executed. In this way, the inverted index library can contain the indexes of all visual media saved in the gallery database. In addition, when the user deletes the data media or attribute information in the gallery database, the gallery APP deletes the visual media or attribute information and notifies the CV service. The search module in the CV service can delete the corresponding indexes in the inverted index library. In this way, the update of the data in the inverted index library is realized. The process of data deletion will not be described in detail here. All in all, the index data in the inverted index library is consistent with the visual media in the gallery data. In this way, an accurate search function can be provided to the gallery APP.

[0235] In the embodiment of the present application, on the one hand, the method not only establishes indexes for incremental images, enabling users to search for matching images in the subsequent search stage, but also establishes indexes for incremental videos, enabling users to search for matching videos in the subsequent search stage, providing users with more and more comprehensive search results and improving the user experience.

[0236] On the second hand, the method performs semantic understanding on the representative frames in incremental images and incremental videos to obtain visual semantic vectors, and constructs indexes based on the visual semantic vectors. In this way, in the subsequent search stage, users can perform fuzzy searches through search statements, and can more accurately search for images and / or videos that meet the actual intentions of users, improving the search fineness and the user experience.

[0237] On the third hand, the method performs analysis of the relationship between characters on the representative frames in incremental images and incremental videos to obtain information about the relationship between characters, and constructs indexes based on the information about the relationship between characters. In this way, in the subsequent search stage, users can search for matching images or videos through the relationship between characters, providing users with more dimensional search methods and improving the user experience.

[0238] On the fourth hand, the method can segment the video based on semantics and then perform relevant processing based on the video segments without manual segmentation by the user. Moreover, since the video is dynamic, there are often many scene changes in a video. By segmenting the video, different scenes are separated, and not only can multiple image semantics and multiple relationships between characters in the video be recognized, but the information obtained from the video is more comprehensive and accurate. Moreover, for each video segment, the interference of irrelevant information can be reduced, and the accuracy of the constructed index can be improved, so that the user's search intention can be accurately located in the subsequent search stage, and the accuracy of the subsequent search can be improved.

[0239] Fifthly, when performing processing such as analyzing the relationships between characters and semantic understanding, the method is based on representative frames extracted from each video segment. In this way, it is not necessary to process each frame image in the video segment separately, which can not only reduce noise information, but also reduce the amount of calculation, storage, etc., thereby reducing resource overhead.

[0240] 2. Search stage

[0241] Exemplarily, Figure 16 is a flowchart of another visual media search method provided by an embodiment of the present application. As Figure 16 shown, the search stage of this method includes:

[0242] S201. In response to the operation of the user inputting a search text, the gallery APP sends the search text to the search module.

[0243] Referring to Figure 7 , for example, the user can enter the search text "a boy standing outdoors" in the search box. After the gallery APP receives the text input by the user, it sends the text to the search module.

[0244] S202. After receiving the search text, the search module sends a language understanding request to the natural language understanding module.

[0245] The language understanding request is used to request natural language understanding of the search text. Optionally, the search text can be carried in the language understanding request.

[0246] S203. In response to the language understanding request, the natural language understanding module performs natural language understanding on the search text, and identifies semantic entities, relationships between characters, category information, etc. included in the search text.

[0247] Optionally, the natural language understanding module can identify semantic entities, relationships between characters, and category information based on a natural language understanding model and / or named entity recognition (NER) technology. For example, when performing natural language understanding on the search text "traveled with family last year", a semantic entity "last year" and a relationship between characters "family" can be obtained. Another example is that when performing semantic entity recognition on the search text "buildings photographed in Beijing during the National Day", the two semantic entities "National Day" and "Beijing" and a category information "buildings" can be obtained.

[0248] S204. The natural language understanding module returns the identified semantic entities, relationships between characters, and category information to the search module.

[0249] S205. The search module sends a text semantic understanding request to the multimodal understanding module.

[0250] A text semantic understanding request is used to request semantic understanding of the search text.

[0251] Optionally, the search text may be carried in the text semantic understanding request.

[0252] S206. In response to the text semantic understanding request, the multimodal understanding module performs semantic understanding on the search text to obtain a text semantic vector.

[0253] Specifically, the multimodal understanding module may include a pre-trained text semantic understanding model. By inputting the search text into the image semantic understanding model, the corresponding semantic vector can be obtained. In this embodiment, the semantic vector corresponding to the text is referred to as the text semantic vector.

[0254] S207. The multimodal understanding module returns the text semantic vector to the search module.

[0255] S208. Based on the semantic entity, person relationship, category information, and text semantic vector, the search module performs multi-way recall in the inverted index library to obtain multi-way recall results. The recall results include images and / or video segments.

[0256] Specifically, the search module may perform semantic entity recall (also known as entity recall) in the inverted index library based on the semantic entity, that is, search for video segments or images matching the semantic entity in the inverted index library to obtain semantic entity recall results; perform person relationship recall in the inverted index library based on the person relationship, that is, search for video segments or images with person relationship information matching the person relationship in the search text in the inverted index library to obtain person relationship recall results; perform category information recall in the inverted index library based on the category information, that is, search for video segments or images with labels matching the category information in the inverted index library to obtain label recall results; perform semantic vector recall (abbreviated as vector recall) in the inverted index library based on the text semantic vector, that is, search for video segments or images with visual semantic vectors matching the text semantic vector in the inverted index library to obtain vector recall results.

[0257] It should be noted that the above "matching" may be consistency or the similarity may meet a preset condition. For example, the similarity is greater than a preset threshold.

[0258] Optionally, if the recalled object is an image, the multi-way recall result may be represented by the ID of the image; if the recalled object is a video segment, the multi-way recall result may be represented by the ID of the video segment and the ID of the video to which it belongs.

[0259] Optionally, when performing recall, the search module can relax the recall conditions to a certain extent, that is, appropriately expand the search scope to make the recall results more comprehensive. For example, when the semantic entity is "National Day", when determining the time range corresponding to the semantic entity, the time range can be appropriately expanded, and September 29th to October 9th can be determined as the time range corresponding to "National Day".

[0260] S209. The search module filters the multi-channel recall results according to the search text to obtain the hit results, and the hit results include hit images and / or hit video segments.

[0261] That is to say, the recall results are filtered to obtain the hit results. The hit results can include images, which are called hit images, and the hit results can also include video segments, which are called hit video segments. In addition, for the convenience of description, the video to which the hit video segment belongs is called the hit video (also called the target video).

[0262] As described above, when performing vector recall, the recall conditions may be relaxed, so the recall results may contain more results that do not match the search statement. In addition, when performing multi-channel recall, each channel of recall may not consider the recall conditions of other channels of recall, so the recall results may also contain more results that do not match the search statement. In view of this, the search module can perform time filtering, space filtering, and person relationship filtering on the multi-channel recall results to improve the accuracy of the final search results.

[0263] Specifically, the search module can filter out the results in the multi-channel recall results that do not match the time included in the search text. For example, if the time in the search text is "last year", the results in the multi-channel recall results whose acquisition time does not overlap with "last year" can be filtered out. For example, the recall result with the acquisition time of January 1st this year can be filtered out.

[0264] The search module can filter out the results in the multi-channel recall results that do not match the location included in the search text. For example, if the location in the search text is "Haidian District, Beijing", the recall result with the acquisition location of "Xicheng District, Beijing" in the multi-channel recall results can be filtered out.

[0265] The search module can filter out the results in the multi-channel recall results that do not match the person relationship included in the search text. For example, if the person relationship in the search text is "me and my son", the recall results with the person relationship information including only "me" and the recall results with the person relationship information including only "son" in the multi-channel recall results can be filtered out.

[0266] S210. The search module determines the matching scores of each hit result, and performs weighted fusion and sorting on the hit results based on the matching scores to obtain the search results.

[0267] The matching score represents the degree of relevance between the hit result and the search text. The sorting result can include a fused sorting result and an in-video sorting result. The fused sorting result is the result obtained by sorting all hit videos and all hit images together as the sorting objects. The in-video sorting result refers to the result obtained by sorting all hit video segments included in a certain hit video.

[0268] Specifically, the search module can perform a matching degree scoring on each hit result from multiple matching dimensions. Then, based on the scoring results of each matching dimension, calculate the information entropy of each matching dimension. Determine the weight of each matching dimension through the information entropy. Then, based on the weights of each matching dimension, calculate the matching scores of each hit image and each hit video segment. After that, for a hit video, use the highest matching score of the hit video segments it contains as the matching score of the hit video. Based on the matching scores, sort all hit videos and hit images to obtain the fused sorting result. In addition, for each hit video segment, based on the matching scores, sort all hit video segments included in the hit video to obtain the intra-segment sorting result.

[0269] This step will be described in detail in the subsequent embodiments.

[0270] S211. The search module returns the search result to the gallery APP.

[0271] The search result takes the information of the hit images (such as ID), the information of the hit video segments (such as the ID of the video to which they belong, start time, end time, etc.), the fused sorting result, and the in-video sorting result of each hit video as the search result and returns it to the gallery APP.

[0272] S212. The gallery APP displays the search result.

[0273] The gallery APP displays the thumbnails of each hit image and hit video according to the fused sorting result in the search result. For a hit video, relevant information can also be displayed in the thumbnail according to the information of the hit video segments it contains, as Figures 6 to 9 shown.

[0274] In addition, when the user performs a play operation on a hit video, the gallery APP responds to the user's click, plays the video, and displays corresponding play effects and matching degree rankings and other information according to the in-video sorting result of the video, as Figure 10 and Figure 11 shown.

[0275] In this embodiment, natural language understanding is performed on the search content input by the user to identify semantic entities, person relationships, category information, etc. therein, and semantic understanding of the search content is performed to obtain a text semantic vector. Then, multi-channel recall is performed based on the semantic entity, person relationship, category information, and text semantic vector. In this way, it is possible to perform all-round search from multiple dimensions, match with the search text from multiple dimensions, greatly improve the matching degree between the search result and the search text, and improve the user experience. Moreover, the hit results are weighted and fused and sorted based on the matching scores to obtain the search results, enabling the user to preferentially see visual media with a high matching degree with the search text, thereby improving the user experience.

[0276] The video segmentation process will be further described below with reference to the accompanying drawings.

[0277] Exemplarily, Figure 17 FIG. is a schematic flowchart of a video segmentation process provided by an embodiment of the present application. As Figure 17 shown, in the above step S104, "performing segmentation processing on the incremental video to obtain multiple video segments and their attribute information, and extracting representative frames from each video segment" includes the following steps. The execution subject of the following steps is the video segmentation module, which will not be elaborated.

[0278] S301. Perform frame splitting processing on the incremental video to obtain multiple frame images.

[0279] S302. Based on the labels and parameters of each frame image, determine the segmentation score of the frame image, where the segmentation score characterizes the degree of change between the frame image and its adjacent frame image.

[0280] It can be understood that when playing a video, the content displayed by the video changes continuously as the frame images are played in sequence, but the degree of change is different. Therefore, similar (i.e., small degree of change) frame images can be classified into one video segment.

[0281] In some embodiments, the segmentation score calculated through the frame label of the frame image and the frame parameter of the frame image can characterize the degree of change between a frame image and its adjacent frame image, and then the video is segmented based on the segmentation score. The frame parameter is used to characterize the display characteristics of the frame image and can also be called an image parameter. Exemplarily, the frame parameter of the frame image may include the jitter degree, clarity, pixel value, etc. of the frame image, and the present application does not limit this.

[0282] Among them, the jitter of the frame image refers to the phenomenon that the content displayed by the frame image jitters or shakes during video playback. Exemplarily, when a user holds a mobile phone to take a picture and moves the mobile phone when hoping to take pictures of another scene, a significant jitter may occur. The jitter degree is used to characterize the jitter degree of the frame image.

[0283] The clarity of a frame image refers to the clarity of the fine patterns and their boundaries in the frame image.

[0284] The pixel values of a frame image can represent the brightness of the frame image.

[0285] In some embodiments, based on the frame label and frame parameters of the frame image, the segmentation score of the frame image can be calculated by combining the following formula (1), and then it can be determined whether to determine it as the start frame or the end frame of a video segment according to the segmentation score of the frame image. Formula (1) is as follows:

[0286] y = α × frame A + β × frame B + γ × frame Y + δ × frame T (1)

[0287] Wherein, y represents the segmentation score of the frame image, frame A represents the jitter degree score of the frame image, frame B represents the clarity change score of the frame image, frame Y represents the label change score of the frame image, frame T represents the pixel change score of the frame image, and α, β, γ, and δ respectively represent frame A , frame B , frame Y and frame T coefficients. In some embodiments, α, β, γ, and δ can be values preset manually according to the influence degrees of the jitter degree, clarity change, label change, and pixel change of the frame image on the segmentation score of the frame image.

[0288] It should be noted that the process of determining the segmentation score of the frame image based on the frame label and multiple frame parameters above is only an example. It can also be based on one or more of the frame label and multiple frame parameters to determine the segmentation score of the frame image, and the present application does not limit this.

[0289] In some embodiments, video jitter detection methods such as the optical flow method, feature point matching method based on image displacement, and based on image gray distribution characteristics can be used to detect the jitter degree of multiple frame images in a video. Then, based on the correspondence between the preset jitter degree range and the jitter degree score, the jitter degree score of the frame image can be determined.

[0290] In some embodiments, a clarity detection tool can be used to determine the clarity of multiple frame images in a video, and then compare the clarity of a frame image with that of an adjacent frame image to determine the clarity change value of the frame image. Subsequently, based on the correspondence between the preset clarity change value range and the clarity change score, the clarity change score of the frame image is determined. Exemplarily, the video quality detection tool can be open-source software such as FFmpeg, Video Quality Measurement Tool, etc.

[0291] In some embodiments, the frame label of a frame image can be compared with the frame label of an adjacent frame image, and the comprehensive label change of the frame image is determined based on the quantity change and content change of the frame label of the frame image. Exemplarily, the union and intersection of the frame label of a frame image and the frame label of its adjacent frame image can be calculated. The quantity change of the frame labels in the union can reflect the quantity change of the frame labels of the frame image, and the quantity change of the frame labels in the intersection can reflect the content change of the frame labels. When there is a change in the quantity of the frame labels in the union or the quantity of the frame labels in the intersection, the label change score of the frame image can be determined according to the correspondence between the preset quantity change range of the frame labels and the label change score.

[0292] In some embodiments, a pixel detection tool can be used to determine the pixels of multiple frame images in a video, and then compare the pixels of a frame image with those of an adjacent frame image to determine the pixel change value of the frame image. Subsequently, based on the correspondence between the preset pixel change value range and the pixel change score, the pixel change score of the frame image is determined. Exemplarily, the pixel detection tool can be plugins such as PixelStick, MeasureIt, Guides, etc.

[0293] S303. Segment the incremental video according to the segment scores of each frame image.

[0294] In some embodiments, the degree of change in the content presented by two adjacent frame images in a video can be positively correlated with the magnitude of the segment score of the frame image, that is, the larger the segment score of the frame image, the more obvious the degree of change in the content it presents compared with the adjacent frame image.

[0295] Exemplarily, a frame image can be compared with the previous frame image, and a segment score threshold is preset in advance. The frame image whose segment score exceeds the segment score threshold is determined as the starting frame of a video segment. As shown in the above formula (1), the higher the jitter degree of the frame image, the higher the corresponding value; the greater the clarity change of the frame image compared with the previous frame image, the higher the corresponding value; the greater the label change of the frame image compared with the previous frame image, the higher the frame A corresponding value; the greater the clarity change of the frame image compared with the previous frame image, the higher the corresponding value; the greater the label change of the frame image compared with the previous frame image, the higher the frame B corresponding value; the greater the label change of the frame image compared with the previous frame image, the higher the frameY The higher the corresponding value; compared with the previous frame image, the greater the change in the label of this frame image, frame Y The higher the corresponding value; compared with the previous frame image, the greater the change in pixels of this frame image, frame - The higher the corresponding value.

[0296] Exemplarily, a frame image can be compared with its next frame image, and a segmentation score threshold can be preset, and the frame image whose segmentation score exceeds the segmentation score threshold is determined as the end frame of a video segment.

[0297] In some other embodiments, the degree of change in the content shown by two adjacent frame images in the video can also be negatively correlated with the size of the segmentation score of the frame image, that is, the smaller the segmentation score of the frame image, the more obvious the degree of change in the content shown compared with the adjacent frame images, that is, the lower the segmentation score of the frame image, the greater the possibility of determining it as the start frame or the end frame of a video segment.

[0298] Exemplarily, a frame image can be compared with its previous frame image, and a segmentation score threshold can be preset, and the frame image whose segmentation score is lower than the segmentation score threshold is determined as the start frame of a video segment. As shown in the above formula (1), the higher the jitter degree of this frame image, frame A The lower the corresponding value; compared with the previous frame image, the greater the change in clarity of this frame image, frame B The lower the corresponding value; compared with the previous frame image, the greater the change in the label of this frame image, frame Y The lower the corresponding value; compared with the previous frame image, the greater the change in the label of this frame image, frame Y The lower the corresponding value; compared with the previous frame image, the greater the change in pixels of this frame image, frame T The lower the corresponding value.

[0299] S304. Determine the corresponding representative frames from each video segment.

[0300] In some embodiments, the representative frame can be the start frame, end frame, middle frame or random frame in a video segment.

[0301] Exemplarily, assuming that a video segment contains 19 frame images, then the first frame image (start frame), the 19th frame image (end frame) or the 10th frame image (middle frame) among these 19 frame images can be used as the representative frame of this video segment according to the time sequence, or a frame image (random frame) can be randomly selected from these 19 frame images as the representative frame of this video segment.

[0302] In some embodiments, the optimal frame can also be determined from a video segment as the representative frame based on preset rules related to the frame parameters and frame labels of the frame images.

[0303] Exemplarily, the score of a frame image can be calculated based on the frame parameters and frame labels of the frame image. The more labels the frame image has, the higher its score; the lower the jitter degree of the frame image, the higher its score; the higher the clarity of the frame image, the higher its score; and the lower the pixel change of the frame image compared to the pixels of the previous frame image, the higher its score. Finally, the frame image with the highest score (i.e., the optimal frame) in a video segment can be determined as the representative frame.

[0304] In this implementation manner, considering various aspects such as the true parameters and frame labels in the video segment, the optimal frame is determined from the video segment as the representative frame, which reduces the probability of selecting an interfering frame as the representative frame, reduces the perturbation of irrelevant information, improves the accuracy of indexing, and thus improves the accuracy of video search.

[0305] Next, in combination with Figure 18 The process of determining the representative frame in the video will be introduced in detail.

[0306] As Figure 18 shown, the video includes 120 frame images. One frame image is compared with the previous frame image, and the respective segment scores corresponding to these 120 frame images are calculated in combination with the above formula (1). The frame image whose segment score exceeds the segment score threshold is used as the starting frame of a video segment.

[0307] In some embodiments, these 120 frame images are divided into 3 video segments, namely video segment 1, video segment 2, and video segment 3, and the middle frame of each video segment is determined as its corresponding representative frame. Video segment 1 includes 40 frame images, video segment 2 includes 51 frame images, and video segment 3 includes 29 frame images. That is, the 20th frame image in video segment 1 is used as the representative frame corresponding to video segment 1, the 26th frame image in video segment 2 is used as the representative frame corresponding to video segment 2, and the 15th frame image in video segment 3 is used as the representative frame corresponding to video segment 3.

[0308] It should be noted that the number of frame images in video segment 1 is 40, which is an even number. Both the 20th frame image and the 21st frame image are middle frames, and any one of the middle frames can be determined as the representative frame corresponding to video segment 1. This application does not make any limitations in this regard.

[0309] The multi-way recall process will be further described below.

[0310] Exemplarily, Figure 19 is a schematic flow diagram of an example multi-way recall process provided by an embodiment of the present application. As Figure 19As shown, the above step "S208. The search module performs multi-way recall in the inverted index library based on the semantic entity, the relationship between people, the category information, and the text semantic vector to obtain the multi-way recall result" includes the following steps. The execution entity of the following steps is the search module, which will not be elaborated further.

[0311] S401. Based on the attribute information of the visual media in the inverted index library, perform the search process corresponding to the semantic entity in the search text to obtain the semantic entity recall result.

[0312] Specifically, obtain the attributes of each visual media among multiple visual media in the inverted index library; match the semantic entity in the search text with the attribute information of each visual media to determine the visual media matched by the semantic entity; use the visual media matched by the semantic entity as the semantic entity recall result. When the number of semantic entities in the search text is one, the semantic entity recall result includes the visual media matched by this one semantic entity; when the number of semantic entities in the search text is multiple, the semantic entity recall result includes the visual media matched by each of these multiple semantic entities.

[0313] In practical applications, different matching methods can be adopted for different semantic entities.

[0314] Exemplarily, when the semantic entity is a semantic entity related to time, determine the time range corresponding to the semantic entity (i.e., the time search range); match the time range with the attribute information of the acquisition time of each visual media to determine the visual media whose acquisition time is within this time range; use the visual media whose acquisition time is within this time range as the visual media matched by the semantic entity. For example: the semantic entity is "National Day", and its corresponding time range is "from October 1st to October 7th". The acquisition time of image 1 is "October 2nd", and the acquisition time of video 2 is "October 8th". Then, according to the above matching method, image 1 matches the semantic entity "National Day", and video 2 does not match the semantic entity "National Day".

[0315] Exemplarily, when the semantic entity is a semantic entity related to location, determine the geographical range corresponding to the semantic entity (i.e., the geographical search range); match the geographical range with the attribute information of the acquisition location of each visual media to determine the visual media whose acquisition location is within this geographical range; use the visual media whose acquisition location is within this geographical range as the visual media matched by the semantic entity. For example: the semantic entity is "Beijing", and its corresponding geographical range is the entire Beijing. The acquisition location of image 3 is "Xicheng District, Beijing", and the acquisition location of video 4 is "Nankai District, Tianjin". Then, according to the above matching method, image 3 matches the semantic entity "Beijing", and video 4 does not match the semantic entity "Beijing".

[0316] Exemplarily, when the semantic entity is a semantic entity related to a person's name, the semantic entity can be textually matched with the names of people in each visual medium, and then the visual medium matched by the semantic entity can be determined. For example: If the semantic entity is "Zhang San", then an image with the person name attribute of "Zhang San" matches this semantic entity.

[0317] S402. Based on the attribute information of the visual media in the inverted index library, perform the search process corresponding to the category information to obtain the label recall result.

[0318] Specifically, the labels in the multiple media attribute information in the inverted index library can be obtained, and the category information is matched with the labels to determine the visual media matched by this category information; the visual media matched by this category information is used as the label recall result.

[0319] S403. Based on the person relationship information of the visual media in the inverted index library, perform the search process corresponding to the person relationship in the search text to obtain the person relationship recall result.

[0320] Exemplarily, the person relationship information corresponding to each visual media can be textually matched with the person relationship identified in the search text, and then the visual media matched by the person relationship can be determined. For example: If the person relationship is "son", then an image or video segment with the person relationship information of "son" matches the person relationship in this search text.

[0321] S404. Based on the visual semantic vectors of the visual media in the inverted index library, perform the vector search process corresponding to the search text to obtain the first vector recall result.

[0322] Specifically, obtain the visual semantic vectors of multiple visual media in the inverted index library; calculate the vector similarity between the text semantic vector of the search text and the visual semantic vectors of each visual media; according to the vector similarity, determine one or more visual media (images or video segments) from multiple visual media whose visual semantic vectors match this text semantic vector; use these one or more visual media as the first vector recall result.

[0323] Optionally, images or video segments with a vector similarity greater than a preset threshold can be selected for recall. Optionally, the multiple obtained vector similarities can also be sorted, and the N indexes with higher vector similarities are selected for label recall. Where N is an integer greater than 0. Exemplarily, N can be the number of preset vector recall results, such as 5, 8, or 10, etc.

[0324] The vector similarity refers to the degree of similarity between two vectors, and can be calculated by various methods. Exemplarily, the similarity degree can be determined by calculating the cosine similarity of the two vectors, or other methods can also be used. This application does not make any limitations in this regard.

[0325] Exemplarily, assume that the multimodal model can encode visual media and search text into a k-dimensional vector space to obtain corresponding visual semantic vectors and text semantic vectors. Assume that the search text is Q, and the corresponding text semantic vector V Q ={α j}, j = 1, 2, …, k. A total of m images (representative frames in the image or video segment) are recalled to form an image set R, and its vector representation is as follows:

[0326]

[0327] Among them, the i-th row corresponds to the visual semantic vector of the i-th image, and the value range of i is [1, m].

[0328] S405. Rewrite the search text.

[0329] Specifically, delete the information in the search text that is irrelevant to the visual content to obtain the rewritten search text. As mentioned above, the information irrelevant to the visual content may include: location, time, name, file attribute, person name, label, person relationship information, etc.

[0330] Exemplarily: for the search text "sky photographed this year", its rewritten search text is "photographed sky".

[0331] In practical applications, after deleting the information in the search text that is irrelevant to the visual content, there may be some redundant stop words. For example: for the search text "sky photographed in Beijing this year", after deleting "this year" and "Beijing", the stop word "in" becomes a redundant word, so it also needs to be deleted. Specifically, delete the semantic subject irrelevant to the visual content in the search text and its related stop words to obtain the rewritten search text. Exemplarily, the rewritten search text corresponding to the search text "sky photographed in Beijing this year" is "photographed sky".

[0332] It should be noted that there is no order restriction for the above steps S401 to S405. In an optional example, to improve efficiency, these 5 steps can be executed simultaneously. The same is true for other steps in the embodiments of the present application, and there is no restriction on the execution order.

[0333] S406. Based on the visual semantic vectors of the visual media in the inverted index library, perform the search corresponding to the rewritten search text to obtain the second vector recall result.

[0334] Specifically, obtain the visual semantic vectors of multiple visual media in the inverted index library; calculate the vector similarity between the second text semantic vector of the rewritten search text and the visual semantic vectors of each visual media; determine one or more visual media whose visual semantic vectors match the second text semantic vector from the multiple visual media according to the vector similarity; and these one or more visual media can be used as the second vector recall result.

[0335] Exemplarily, the rewritten search text is Q′, and the corresponding vector V Q′ ={α′ j}, j = 1, 2, …, k, and a total of m′ images (representative frames in the image or video segment) are recalled to form an image set R′, and its vector representation is:

[0336]

[0337] Among them, the i-th row corresponds to the visual semantic vector of the i-th image, and the value range of i is [1, m′].

[0338] S407. Calculate the semantic proportion of the visual content in the search text according to the first vector recall result and the second vector recall result.

[0339] The following will introduce a method for determining the semantic proportion of visual content:

[0340] 1) Determine the first representative visual semantic vector V I corresponding to the first vector recall result and the second representative visual semantic vector V I′ .

[0341] 2) Determine the difference between the text semantic vector V Q and the first representative visual semantic vector V I to obtain the first difference vector (V Q - V I ).

[0342] 3) Determine the difference between the rewritten text semantic vector V Q′ and the second representative visual semantic vector V I′ to obtain the second difference vector (V Q′ - V I′ ).

[0343] 4) Determine the semantic proportion of the visual content in the search text according to the vector similarity between the first difference vector (V Q - V I ) and the second difference vector (V Q′ - V I′ ).

[0344] Among them, the semantic proportion of visual content is positively correlated with the vector similarity.

[0345] In the above step 1), in one example, the vector similarity between the visual semantic vector of each visual medium in the first vector recall result and the text semantic vector of the search text can be obtained; the visual media in the first vector recall result are sorted according to the vector similarity from high to low; the average vector of the visual semantic vectors of the top T (T≥1) visual media in the sorting is used as the first representative visual semantic vector V I .

[0346] Continuing with the above example, the first representative visual semantic vector V I is:

[0347]

[0348] The vector similarity between the visual semantic vector of each visual medium in the second vector recall result and the text semantic vector of the rewritten search text can be obtained; the visual media in the second vector recall result are sorted according to the vector similarity from high to low; the average vector of the visual semantic vectors of the top H (H≥1) visual media in the sorting is used as the second representative visual semantic vector V I′ .

[0349] Continuing with the above example: the second representative visual semantic vector V I′ is:

[0350]

[0351] The values of H and T above can be the same or different, and the embodiments of the present application do not make specific limitations on this. In another example, the visual media in the first vector recall result can be clustered according to the vector similarity between the visual semantic vector of each visual medium in the first vector recall result and the text semantic vector of the search text, and the clustering center point can be obtained; the visual semantic vector of the clustering center point is used as the first representative visual semantic vector.

[0352] The visual media in the second vector recall result can be clustered according to the vector similarity between the visual semantic vector of each visual medium in the second vector recall result and the second text semantic vector of the rewritten search text, and the clustering center point can be obtained; the visual semantic vector of the clustering center point is used as the second representative visual semantic vector.

[0353] In the embodiments of the present application, the representative visual semantic vector is used to represent the visual semantics of the entire vector recall result. The specific clustering algorithm can be selected according to actual needs, and the embodiments of the present application do not make any limitations on this.

[0354] In the above (4), in an optional embodiment, the first difference vector (V Q -V I ) and the second difference vector (V Q′ -V I′ ) can be directly used as the semantic proportion of the visual content in the search text for the vector similarity between them.

[0355] Among them, the greater the vector similarity between the first difference vector (V Q -V I ) and the second difference vector (V Q′ -V I′ ), the greater the semantic proportion of the visual content in the search text; conversely, the smaller the semantic proportion of the visual content in the search text.

[0356] It should be noted that in the embodiments of the present application, the calculation method of the semantic proportion of the visual content in the search text draws on the word analogy characteristics of the distribution representation vector (Distribution Representation), that is, the additivity of word meanings is directly reflected in the additivity of the distribution representation vector.

[0357] S408. If the semantic proportion of the visual content in the search text is less than the preset proportion threshold, discard the vector recall result, and use the semantic entity recall result, label recall result, and person relationship recall result as the final multi-way recall result.

[0358] Generally, when the semantic proportion of the visual content in the search text is less than the preset proportion threshold, it means that the search text does not contain any semantic entities related to visual content, but only contains information unrelated to visual content, such as semantic entities related to time and semantic entities related to space. Therefore, the vector recall result can be discarded, which can reduce the noise information in the search results and improve the search accuracy.

[0359] S409. If the semantic proportion of the visual content in the search text is greater than or equal to the preset proportion threshold, determine the vector recall result as the vector recall result, and use the vector recall result, semantic entity recall result, label recall result, and person relationship recall result as the final multi-way recall result.

[0360] When the semantic proportion of the visual content in the search text is greater than or equal to the preset proportion threshold, it means that the search text contains semantic entities related to visual content. Therefore, the vector recall result needs to be included in the final recall result to improve the search accuracy and comprehensiveness.

[0361] Next, with reference to the accompanying drawings, the weighted fusion sorting process of the hit video and hit image will be further described.

[0362] Exemplarily,Figure 20 This is a schematic flowchart of a weighted fusion sorting process provided by an embodiment of the present application. As Figure 20 shown, in the above step S210, "the search module determines the matching scores of each hit result, and performs weighted fusion sorting on the hit results based on the matching scores to obtain the search result" includes the following steps. The execution subject of the following steps is the search module, which will not be elaborated.

[0363] S501. Determine the matching dimension.

[0364] The matching dimension represents the angle at which the hit result is evaluated for the matching degree with the search text. Optionally, the matching dimension can be a fixed dimension. For example, acquisition time, acquisition location, label, vector, person relationship information, etc.

[0365] In one embodiment, the semantic vector can be used as a matching dimension. At the same time, other matching dimensions can be determined according to the search text. Specifically, the dimension corresponding to the result (abbreviated as the recognition result) identified from the search text through natural language understanding can be determined as the matching dimension. If the recognition result includes a semantic subject related to time, the acquisition time is used as the matching dimension; if the recognition result includes a semantic subject related to location, the acquisition location is used as the matching dimension. If the recognition result includes category information, the label is used as the matching dimension. If the recognition result includes person relationship, the person relationship information is used as the matching dimension. That is, whichever dimension's corresponding information is included in the search text, that dimension is used as the matching dimension.

[0366] S502. Score the matching degree of each hit result in each matching dimension to obtain the corresponding dimension score.

[0367] The dimension score corresponding to a certain matching dimension represents the matching degree of the hit result and the search text in this dimension. For example, the dimension score corresponding to the vector dimension of Image 1 represents the matching degree of Image 1 and the search text in the vector dimension. Optionally, the dimension score corresponding to the vector dimension can be obtained through vector similarity calculation. The dimension score corresponding to the vector dimension is positively correlated with the vector similarity.

[0368] Exemplarily, taking the hit results including Image 1, video segments 2.1 and 2.2 of Video 2, and Image 3, and the matching dimensions including Dimension A, Dimension B, and Dimension C as an example, score each hit result in each matching dimension. The scoring result can be shown in Table 1, for example. Among them, video segment 2.1 represents video segment 1 in Video 2, and video 2.2 represents video segment 2 in Video 2.

[0369] Table 2

[0370] Dimension A Dimension B Dimension C Image 1 0.30 0.35 0.55 Video Segment 2.1 0.40 0.38 0.42 Video Segment 2.2 0.42 0.35 0.41 Image 3 0.38 0.37 0.45

[0371] It should be noted that Table 2 is only for illustration, does not represent actual data, and does not impose any limitation on this application.

[0372] S503. Determine the matching scores of each hit result based on the dimensional scores of each hit result in each matching dimension.

[0373] The matching score is obtained by calculating the dimensional scores, and it is a comprehensive result representing the matching degree of the hit result and the search result in all dimensions. Therefore, the matching score is also called the comprehensive score or the comprehensive matching score, etc.

[0374] There are various calculation methods for the matching score, which will be described in detail in the subsequent embodiments.

[0375] S504. Sort the hit images and hit videos according to the matching scores to obtain a fusion sorting result; sort the hit video segments in each hit video to obtain an in-video sorting result. Among them, the highest matching score in the hit video is used as the matching score of the hit video.

[0376] There are various ways of fusion sorting and in-video sorting. In a specific embodiment, the sorting can be achieved according to the following process: 1) Sort all hit results in descending order of the matching scores to obtain set A. 2) Write the hit results in set A into sorting set B in sequence. When writing a hit result into set B, if the hit result is a hit image, it is directly written after the existing elements in set B; if the hit result is a hit video segment, it is written into the subset corresponding to the hit video to which the hit video segment belongs. In this way, in the finally obtained sorting set B, each hit image and the first hit video segment of each subset are arranged in descending order of the matching scores, and the hit video segments in each subset are arranged in descending order of the matching scores.

[0377] For example, multiple hit results are arranged in descending order of the matching scores to obtain set A, and set A = {Image 2, Video Segment 1.2, Image 1, Video Segment 2.1, Video Segment 1.1, Video Segment 2.2}. Then, the sorting is achieved according to the following process:

[0378] 1) Establish an empty sorting set B;

[0379] 2) Write the image 2 with the highest matching score in set A into the first position in set B to obtain set B = {Image 2}.

[0380] 3) The element with the second highest matching score in set A is video segment 1.2. Therefore, establish subset 1 corresponding to the video (Video 1) to which video segment 1.2 belongs at the position after image 2 in set B, and write video segment 1.2 into the first position in subset 1 to obtain set B = {Image 2, [Video Segment 1.2]}.

[0381] 4) Write the image 1 with the third-highest matching score in set A to the position after subset 1 in set B, obtaining set B = {image 2, [video segment 1.2], image 1}.

[0382] 5) The element with the fourth-highest matching score in set A is video segment 2.1. Therefore, create subset 2 corresponding to the video (video 2) to which video segment 2.1 belongs at the position after image 1 in set B, and write video segment 2.1 to the first position in subset 2, obtaining set B = {image 2, [video segment 1.2], image 1, [video segment 2.1]}.

[0383] 6) The element with the fifth-highest matching score in set A is video segment 1.1, and subset 1 corresponding to video 1 to which video segment 1.1 belongs already exists. Therefore, write video segment 1.1 to the position after video segment 1.2 in subset 1, obtaining set B = {image 2, [video segment 1.2, video segment 1.1], image 1, [video segment 2.1]}.

[0384] 7) The element with the sixth-highest (lowest) matching score in set A is video segment 2.2, and subset 2 corresponding to video 2 to which video segment 2.2 belongs already exists. Therefore, write video segment 2.2 to the position after video segment 2.1 in subset 2, obtaining set B = {image 2, [video segment 1.2, video segment 1.1], image 1, [video segment 2.1, video segment 2.2]}. The sorting is completed to obtain the sorting result.

[0385] As can be seen from the finally obtained set B, the matching scores of image 2, video segment 1.2, image 1, and video segment 2.1 decrease in sequence. That is to say, set B contains the integrated sorting result of the hit images and hit videos: image 1, video 1, image 1, video 2. In addition, the matching scores of video segment 1.2 and video segment 1.1 in subset 1 of set B decrease in sequence, and the matching scores of video segment 2.1 and video segment 2.2 in subset 2 of set B decrease in sequence. That is to say, set B contains the intra-video sorting result of each hit video. Thus, it can be seen that according to the above method, traversing the hit results once can obtain both the integrated sorting result and the intra-video sorting result at the same time, simplifying the sorting process and improving the sorting efficiency.

[0386] Optionally, the above set B can be sent as the sorting result to the gallery APP. If the gallery APP needs to display the search result interface shown above Figures 6 to 9 it can determine the integrated sorting result of the hit images and hit videos from set B. If the gallery APP needs to display the interface shown above Figure 11 it can determine the element sorting of the corresponding subset of the hit video from set B to obtain the intra-video sorting result.

[0387] The specific implementation process of the above step "S503. Determine the matching score of each hit result based on the dimensional scores of each hit result in each matching dimension" will be described below.

[0388] In some alternative embodiments, any one of the following three methods can be used to determine the matching score of the hit result:

[0389] Method 1: Sum the dimensional scores of the hit result in all matching dimensions to obtain the matching score of the hit result.

[0390] Method 2: Perform a weighted sum of the dimensional scores of the hit result in all matching dimensions to obtain the matching score of the hit result. Among them, the weights corresponding to each matching dimension can be configured by the user independently.

[0391] Method 3: Input the matching dimensional scores of each hit result in all dimensions into a machine learning model to obtain the matching score of the hit result.

[0392] For the matching scores calculated by using the above Method 1, it is more likely that the matching scores of different hit results are the same or similar. For example, in Table 2 above, the matching scores of Image 1, Video Segment 2.1, and Image 3 are all equal to 1.2, while the matching score of Video Segment 2.2 is 1.18, which is also relatively close to other matching scores. This will result in the inability to sort the hit results or inaccurate sorting results.

[0393] In the above Method 2, when a relatively large number of types of results are recognized through natural language understanding of the search text, that is, when there are more matching dimensions, it is difficult to accurately configure the weights of each matching dimension.

[0394] In the above Method 3, the machine learning model needs to be trained based on a data set. The construction of the data set is inseparable from user data. To obtain a data set with a sufficient sample size, user data needs to be reported to the cloud. However, user data belongs to user privacy content, and users do not want their data to be reported to the cloud side. Therefore, determining the matching score through this method does not meet the "minimization" principle of user data.

[0395] To solve the above problems, in the embodiments of the present application, the weights of each matching dimension are determined by means of information entropy. Information entropy characterizes the degree of variation of information in each matching dimension. In this way, the weights can be determined quantitatively, the accuracy of weight allocation can be improved, thereby improving the accuracy of matching score calculation and the accuracy of sorting.

[0396] Specifically, referring to Figure 21 , in this embodiment, the following steps can be used to determine the matching score:

[0397] S5031. Normalize the dimensional score matrix to obtain a normalized dimensional score matrix (abbreviated as the normalized matrix); the dimensional score matrix includes the dimensional scores of each hit result on each matching dimension.

[0398] In other words, the dimensional score matrix is a matrix composed of the dimensional scores of each hit result on each dimension.

[0399] Suppose there are m hit results and n matching dimensions. The dimensional scores of m hit results on n matching dimensions form a dimensional score matrix X:

[0400] X = (x ij ) m*n , i = 1, 2, …, m; j = 1, 2, …, n (6)

[0401] where x ij is the dimensional score of the i-th hit result on the j-th matching dimension.

[0402] Continuing with the example shown in Table 2 above, the hit results include Image 1, Video Segment 2.1, Video Segment 2.2, and Image 3, and the multiple matching dimensions include Dimension A, Dimension B, and Dimension C. That is, m is 4 and n is 3.

[0403] Normalize the dimensional score matrix X = (x ij ) m*n to obtain the normalized matrix R as:

[0404] R = (r ij ) m*n , i = 1, 2, …, m; j = 1, 2, …, n (7)

[0405]

[0406] where (r ij ) m*n represents the value after normalizing the dimensional score of the i-th hit result on the j-th matching dimension (referred to as the normalized score). That is, the normalized matrix R includes m * n normalized scores. max(x j ) represents the maximum value among the dimensional scores of m hit results on the j-th matching dimension; min(x j ) represents the minimum value among the dimensional scores of m hit results on the j-th matching dimension.

[0407] In practical applications, other standardization or normalization methods can also be used, and the embodiments of this application do not make specific limitations in this regard.

[0408] Exemplarily, the data in the normalized matrix obtained after normalizing the example data in Table 2 above is shown in Table 3.

[0409] Table 3

[0410] Dimension A Dimension B Dimension C Image 1 0.00 0.00 1.00 Video Segment 2.1 0.83 1.00 0.07 Video Segment 2.2 1.00 0.00 0.00 Image 3 0.67 0.67 0.29

[0411] S5032. Calculate the information entropy corresponding to each matching dimension based on the normalized matrix.

[0412] Optionally, formulas (9) and (10) can be used to calculate the information entropy:

[0413]

[0414]

[0415] where e j represents the information entropy corresponding to the j-th matching dimension.

[0416] Since the domain of the ln(x) function is x > 0, optionally, in actual calculation, ln(p ij ) in formula (9) can be replaced by ln(p ij +α), where α << 0.001. This can avoid p ij in ln(p ij ) from taking the value of 0 and improve the distinguishability of the result.

[0417] Exemplarily, for the example data in Table 2 above, the information entropy calculated according to formulas (9) and (10) above is shown in Table 4:

[0418] Table 4

[0419] Dimension A Dimension B Dimension C Information Entropy 1.09 0.67 0.71

[0420] From the above information entropy calculation formula and calculation results, it can be seen that the information entropy is inversely proportional to the degree of variation of the values corresponding to the matching dimension. That is, the smaller the information entropy corresponding to the matching dimension, the greater the degree of variation of the matching degrees of multiple hit results in this matching dimension, and the greater the amount of information provided. It can be considered that the role played by this matching dimension in the comprehensive evaluation of the matching degree is also greater.

[0421] S5033. Calculate the weights corresponding to each matching dimension according to the information entropy.

[0422] Optionally, formula (11) can be used to calculate the initial weights corresponding to each matching dimension:

[0423]

[0424] where d jDenote the initial weight corresponding to the j-th matching dimension.

[0425] It can be seen that the above formula (11) is a monotonically decreasing function of the information entropy. In this embodiment, the smaller the information entropy, the greater the degree of variation; the greater the degree of variation, the greater the initial weight. That is to say, the initial weight is negatively correlated with the information entropy.

[0426] To ensure that the sum of the weights corresponding to multiple matching dimensions is 1 or close to 1, the following calculation can be performed based on the initial weight using formula (12) to obtain the final weight w j :

[0427]

[0428] where w j Denotes the final weight (abbreviation: weight) corresponding to the j-th matching dimension.

[0429] In an alternative embodiment, other monotonically decreasing functions can also be used to calculate the above initial weight, for example This application embodiment does not make specific limitations on this, as long as the initial weight is negatively correlated with the information entropy.

[0430] Exemplarily, for the example data in Table 4 above, the weights of different matching dimensions can be calculated according to the above formulas (11) and (12), as shown in Table 5:

[0431] Table 5

[0432] Dimension A Dimension B Dimension C Weight 0.01 0.52 0.47

[0433] S5034. Based on the weights of each matching dimension and the normalization matrix, perform a weighted sum of the normalized scores of each hit result on each matching dimension to obtain the matching score of each hit result.

[0434] Perform a weighted sum of the normalized scores of any hit result to obtain the matching score of this hit result, which can be specifically calculated using formula (13):

[0435]

[0436] where s i Denotes the matching score of the i-th hit result.

[0437] Exemplarily, according to the weights shown in Table 5 above, for the example data in Table 2 above, the matching scores of each hit result are calculated using the above formula (13) as shown in Table 6:

[0438] Table 6

[0439] Image 1 Video Segment 2.1 Video Segment 2.2 Image 3 Matching Score 0.47 0.56 0.01 0.49

[0440] For the matching score calculation method provided in this embodiment, on the one hand, the matching score is calculated based on weighted summation, which fully considers the contributions of information in different matching dimensions to the final matching degree. Therefore, the obtained matching score is more in line with the actual situation, has higher accuracy, and the probability that the matching scores of different hit results are the same is relatively small, reducing the occurrence of situations where the matching degree cannot be distinguished among hit results. On the other hand, when performing weighted summation, the information entropy of each matching dimension is calculated, and the weight is determined based on the information entropy. There is no need to manually set the weight. Moreover, the information entropy can quantitatively measure the degree of data variation in different matching dimensions and reflect the contribution degree of data in different matching dimensions to the final matching score. Therefore, it can accurately allocate the weights of different matching dimensions and improve the accuracy of matching score calculation. On the third hand, this method does not require constructing a data set, so it does not need to collect user data, can better protect the privacy and security of users, meet the "minimization" principle of user data, and improve the user experience. On the fourth hand, this method performs normalization processing on the dimension score matrix and calculates the information entropy and weight based on the obtained normalization matrix, which can ensure that the matching degrees of hit results in different matching dimensions have a unified dimension and improve the accuracy of weight and matching score calculation.

[0441] The above text details the examples of the visual media search method provided in the embodiments of this application. It can be understood that in order for the electronic device to implement the above functions, it includes the corresponding hardware and / or software modules for executing each function. Those skilled in the art should easily realize that, combined with the units and algorithm steps of each example described in the embodiments disclosed in this article, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of this application.

[0442] The embodiments of this application can divide the electronic device into functional modules according to the above method examples. For example, each function can be corresponding to each functional module, such as a detection unit, a processing unit, a display unit, etc., or two or more functions can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of this application is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0443] It should be noted that all relevant contents of each step involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be repeated here.

[0444] The electronic device provided in this embodiment is used to execute the above-mentioned visual media search method, so the same effects as the above implementation method can be achieved.

[0445] In the case of adopting an integrated unit, the electronic device may further include a processing module, a storage module, and a communication module. Among them, the processing module can be used to control and manage the operations of the electronic device. The storage module can be used to support the electronic device to execute stored program codes and data, etc. The communication module can be used to support the communication between the electronic device and other devices.

[0446] Among them, the processing module can be a processor or a controller. It can implement or execute various exemplary logical blocks, modules, and circuits described in combination with the disclosure of this application. The processor can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, and so on. The storage module can be a memory. The communication module can specifically be a device that interacts with other electronic devices, such as a radio frequency circuit, a Bluetooth chip, a Wi-Fi chip, etc.

[0447] In one embodiment, when the processing module is a processor and the storage module is a memory, the electronic device involved in this embodiment can be a device with Figure 3 the structure shown.

[0448] The embodiment of the present application also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the processor is enabled to execute the visual media search method of any one of the above embodiments.

[0449] The embodiment of the present application also provides a computer program product. When the computer program product runs on a computer, the computer is enabled to execute the above-related steps to implement the visual media search method in the above embodiment.

[0450] In addition, the embodiment of the present application also provides a device. This device can specifically be a chip, a component, or a module. The device may include a processor and a memory connected to each other; among them, the memory is used to store computer execution instructions. When the device runs, the processor can execute the computer execution instructions stored in the memory, so that the chip executes the visual media search method in each of the above method embodiments.

[0451] Among them, the electronic device, the computer-readable storage medium, the computer program product, or the chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved by them can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.

[0452] From the description of the above embodiments, those skilled in the art can understand that, for the convenience and conciseness of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0453] In several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0454] The unit described as a separated component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it can be located in one place, or it can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0455] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0456] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: USB flash drive, mobile hard disk, read only memory (ROM), random access memory (RAM), magnetic disk or optical disk and other various media that can store program codes.

[0457] The above content is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A visual media search method, which is executed by an electronic device, characterized in that, the method includes: displaying a first interface of a gallery application; the first interface includes a search box; receiving a first operation in which a user inputs a first text through the search box; in response to the first operation, searching for a plurality of hit results that match the first text, and the plurality of hit results include at least one hit image and at least one hit video segment, and the at least one hit video segment is a video segment in at least one target video; respectively determining a matching score for each of the hit results, and the matching score characterizes the degree of matching between the hit result and the first text; sorting the at least one hit image and the at least one target video according to the matching scores of each of the hit results to obtain a first sorting result; displaying the at least one hit image and the at least one target video according to the first sorting result.

2. The method according to claim 1, characterized in that, the sorting the at least one hit image and the at least one target video according to the matching scores of each of the hit results to obtain a first sorting result includes: for a first target video, taking the highest score among the matching scores of all first hit video segments as the matching score of the first target video; the first target video is any one of the at least one target video, and the first hit video segment is a hit video segment in the first target video; sorting the at least one hit image and the at least one target video according to the matching scores to obtain the first sorting result.

3. The method according to claim 2, characterized in that, the number of the first hit video segments is multiple, and the method further includes: sorting the multiple first hit video segments according to the matching scores to obtain a second sorting result.

4. The method according to claim 3, characterized in that, after the displaying the at least one hit image and the at least one target video according to the first sorting result, the method further includes: receiving a second operation in which a user plays the first target video; in response to the second operation, sequentially playing the multiple first hit video segments according to the second sorting result.

5. The method according to any one of claims 2 to 4, characterized in that, the displaying the at least one hit image and the at least one target video includes: displaying thumbnails of each of the hit images and thumbnails of each of the target videos; wherein, the thumbnail of the first target video is a thumbnail of a frame image in a second hit video segment, and the second hit video segment is the video segment with the highest matching score among all the first hit video segments; one or more of the following data are displayed in the thumbnail of the first target video: the start playing moment of the second hit video segment, the end playing moment of the second hit video segment, the middle playing moment of the second hit video segment, the total number of the first hit video segments, and the total number of video segments included in the first target video.

6. The method according to any one of claims 1 to 5, characterized in that, the separately determining the matching scores of the respective hit results includes: determining a plurality of matching dimensions according to the first text; determining the dimension scores of the respective hit results on the respective matching dimensions, wherein the dimension score of the first hit result on the first matching dimension represents the matching degree between the first hit result and the first text on the first matching dimension; the first hit result is any one of the at least one hit result, and the first matching dimension is any one of the plurality of matching dimensions; calculating the weights of the respective matching dimensions according to the dimension scores of the respective hit results on the respective matching dimensions; weighted summing the dimension scores of the first hit result on the respective matching dimensions according to the weights of the respective matching dimensions to obtain the matching score of the first hit result.

7. The method according to claim 6, characterized in that, the calculating the weights of the respective matching dimensions according to the dimension scores of the respective hit results on the respective matching dimensions includes: performing a normalization process on a dimension score matrix to obtain a normalized matrix, where the dimension score matrix is a matrix composed of the dimension scores of the respective hit results on the respective matching dimensions; calculating the information entropy of the respective matching dimensions based on the normalized matrix; calculating the weights of the respective matching dimensions according to the information entropy of the respective matching dimensions, wherein the weight of the first matching dimension is negatively correlated with the information entropy of the first matching dimension.

8. A visual media search method, which is executed by an electronic device, characterized in that, the method includes: displaying a first interface of a gallery application; the first interface includes a search box, the search box includes a first text, and the first interface further includes a first thumbnail of a first target video; a first playback moment is displayed in the first thumbnail, and the first playback moment corresponds to a first target frame image in the first target video; in response to a trigger operation of the user on the first thumbnail, playing the first target video starting from the first playback moment; displaying a playback progress bar; the playback progress bar displays first marking information marking the first playback moment and second marking information marking a second playback moment, and the second playback moment corresponds to a second target frame image in the first target video; wherein, both the first target frame image and the second target frame image match the first text.

9. The method according to claim 8, characterized in that, the method further includes: in response to an operation of the user on the second marking information, playing the first target video starting from the second playback moment.

10. The method according to claim 8 or 9, characterized in that, the first target video further includes at least one third target frame image, the third target frame image is different from the first target frame image and the second target frame image, and the third target frame image matches the first text; The first playback time is the earliest one among the first playback time, the second playback time, and the playback times corresponding to each of the third target frame images; Or, The first target frame image is the one with the highest degree of matching with the first text among the first target frame image, the second target frame image, and the at least one third target frame image.

11. The method according to claim 10, wherein, Sorting information of the first target frame image, the second target frame image, and the at least one third target frame image is further displayed in the playback progress bar, and the sorting information represents the sorting of the degree of matching of the frame images with the first text.

12. The method according to claim 11, wherein, The sorting information is represented by one or more of color, text, graphics, and symbols.

13. A visual media search method, which is executed by an electronic device, wherein, The method includes: Displaying a first interface of a gallery application; the first interface includes a search box, the search box includes a first text, and the first interface further includes a first thumbnail of a first target video; a first playback time is displayed in the first thumbnail; In response to a triggering operation of the user on the first thumbnail, playing the first target video starting from the first playback time; Displaying a playback progress bar; marking information of a first playback time period and marking information of a second playback time period are displayed in the playback progress bar, the first playback time period corresponds to a first video segment in the first target video, the first playback time period includes the first playback time; the second playback time period corresponds to a second video segment in the first target video; wherein, there is at least one frame image in the first video segment that matches the first text, and there is at least one frame image in the second video segment that matches the first text.

14. The method according to claim 13, wherein, The playing the first target video starting from the first playback time includes: Playing the first video segment starting from the first frame image of the first video segment; Or, playing the first video segment starting from a first target frame image in the first video segment, and the first target frame image matches the first text.

15. The method according to claim 14, wherein, The method further includes: After the first video segment is played, playing the second video segment.

16. The method according to any one of claims 13 to 15, wherein, First marking information corresponding to the first video segment and second marking information corresponding to the second video segment are further displayed in the playback progress bar; the method further includes: In response to an operation of the user on the second marking information, playing the second video segment.

17. The method according to claim 15 or 16, wherein, The playing the second video segment includes: Playing the second video segment starting from the first frame image of the second video segment; Alternatively, start playing the second video segment from the second target frame image in the second video segment, where the second target frame image matches the first text.

18. The method according to any one of claims 13 to 17, wherein, sorting information of the first video segment and the second video segment is further displayed in the playback progress bar, and the sorting information represents the sorting of the degree of matching of the video segments with the first text.

19. The method according to claim 18, wherein, the sorting information is represented by one or more of color, text, graphics, and symbols.

20. An electronic device, wherein, comprising a memory, a processor, and a display screen; the display screen is used to display the user interface of the gallery; the memory is coupled to the processor, the memory is used to store computer program code, the computer program code includes computer instructions, and the processor calls the computer instructions to cause the electronic device to execute the method according to any one of claims 1 to 19.

21. A computer-readable storage medium, wherein, a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the processor calls instructions to cause the electronic device to execute the method according to any one of claims 1 to 19.

Citation Information

Cited By

  • Visual media search method, electronic device and storage medium

    WO2025103081A1