Visual media searching method and electronic equipment

By introducing character relationship search function in image and video search, the problem of insufficient search experience in the existing technology is solved, richer and more accurate search results are achieved, and user experience is improved.

CN120067373APending Publication Date: 2025-05-30HONOR DEVICE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311580872.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, the experience brought by the image search function to the user needs to be improved, and the character relationship information in the video cannot be effectively utilized, which limits the user's search dimension.

Method used

Provides a visual media search method that allows users to search fuzzyly based on search statements and search for matching images or videos through character relationships. The method uses the method to identify the character relationship information in the search text, searches in the stored visual media of the electronic device, and displays videos and images that match the character relationship.

Benefits of technology

It improves the user experience and provides more dimensions of search methods, allowing users to find visual media related to search statements more conveniently, and enhances the accuracy and comprehensiveness of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067373A_ABST
    Figure CN120067373A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a visual media search method and electronic equipment, and the method comprises the steps: displaying a first interface which comprises a search box; receiving a first operation of inputting a search text in the search box by a user; the search text comprises first information representing a character relationship; in response to the first operation, displaying a first search result; the first search result corresponds to a first video, the first video comprises a first character, the relationship between the first character and a central character is consistent with the first information, and the central character refers to a character in a social center among characters contained in a visual media stored in the electronic equipment. According to the method, video searching can be achieved, searching can be conducted based on the character relation, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic technologies, and in particular, to a visual media search method and an electronic device. Background Art

[0002] With the rapid development of the camera configurations of electronic devices such as mobile phones and tablet computers, taking pictures and shooting videos with electronic devices has become one of the important ways for people to record their lives. Moreover, with the rapid development of aspects such as the storage capacity of electronic devices and network technologies, the images and videos stored in users' electronic devices are also increasing.

[0003] In order to facilitate users to manage and view the images in electronic devices, an image search function is configured in the gallery application or other similar applications of electronic devices. For example, the gallery application in an electronic device can classify the images in the electronic device according to information such as the time and location of image shooting, and users can view relevant images by searching for information such as time and location.

[0004] However, in related technologies, the experience brought by image search to users still needs to be improved. Summary of the Invention

[0005] This application provides a visual media search method and an electronic device. Users can perform fuzzy search based on search statements, and users can search for matching images or videos through the relationship between people, providing users with more dimensions of search methods and improving the user experience.

[0006] In a first aspect, this application provides a visual media search method, which is executed by an electronic device. The method includes: displaying a first interface, where the first interface includes a search box; receiving a first operation in which a user inputs search text in the search box; the search text includes first information representing the relationship between people; in response to the first operation, displaying a first search result; the first search result corresponds to a first video, and the first video includes a first person, and the relationship between the first person and the central person is consistent with the first information, and the central person refers to the person in the visual media stored in the electronic device who is in the social center.

[0007] For the visual media search method provided in the first aspect, when the first information representing the relationship between people is included in the search text input by the user, the electronic device can respond to the user's operation and display the first video whose included relationship between people matches the first information. In this way, on the one hand, the search results include videos, providing users with more and more comprehensive search results and improving the user experience. On the other hand, users can search for matching videos through the relationship between people, providing users with more dimensions of search methods and improving the user experience.

[0008] In one possible implementation, the method further includes: in response to a first operation, displaying second search results, where the second search results correspond to a first image; the first image includes a second person, and the relationship between the second person and the central person is consistent with the first information.

[0009] That is to say, the search results include not only videos but also images. In this way, more and more comprehensive search results are provided to the user, improving the user experience.

[0010] In one possible implementation, the first video includes a plurality of second video segments, and at least one hit video segment is included in the plurality of second video segments. The hit video segment is a video segment that includes a first person among the plurality of second video segments.

[0011] In this implementation, the video is segmented to obtain a plurality of video segments. Since the video is dynamic, there are often many scene changes in a video. By segmenting the video, different scenes are separated. Not only can multiple image semantics and multiple person relationships in the video be recognized, but the information obtained from the video is more comprehensive and accurate. Moreover, for each video segment, the interference of irrelevant information can be reduced, improving the accuracy of the search.

[0012] In one possible implementation, the method further includes: receiving a second operation in which the user clicks on the first search result; in response to the second operation, starting to play the first video from the first frame image of the first hit video segment, or starting to play the first video from the representative frame image; the first hit video segment is one of the at least one hit video segment, and the representative frame image is a frame image in the first hit video segment, and the representative frame image includes the first person.

[0013] In this implementation, when the user clicks on the first search result, that is, when the user clicks on the first video, in response to the user's click, the first video starts to be played from a certain frame of the hit video segment. In this way, the first video can be directly played from the position related to the search text, and the user can directly see the hit video segment related to the search text without waiting for the playback or dragging the playback progress bar to find the hit video segment, improving the user experience.

[0014] In one possible implementation, displaying the first search result includes: displaying a thumbnail of the cover frame image; the cover frame image is a frame image in the first video, and the first playback moment is displayed in the thumbnail, and the first playback moment is within the closed interval formed by the start playback moment and the end playback moment of the first hit video segment.

[0015] In this implementation, the first playback moment is displayed in the thumbnail of the cover frame image, and the user can intuitively know the position of the hit video segment in the first video, improving the user experience.

[0016] In a possible implementation, the first playback moment is: the start playback moment of the first hit video segment, the end playback moment of the first hit video segment, the middle playback moment of the first hit video segment, or the playback moment corresponding to the representative frame image.

[0017] In a possible implementation, the cover frame image is the first frame image of the first video, or the first frame image of the first hit video segment, or the representative frame image.

[0018] In this implementation, the cover frame is the first frame image or the representative frame image of the first hit video segment. Through the cover frame, the user can intuitively learn about the video scenes related to the search text, improving the user experience.

[0019] In a second aspect, the present application provides a visual media search method, which is executed by an electronic device. The method includes: displaying a first interface, where the first interface includes a search box; receiving a first operation in which the user enters a search text in the search box; in response to the first operation, identifying first information representing a person relationship in the search text; according to the first information, searching for a first video among multiple videos; the first video includes multiple second video segments, and at least one hit video segment is included in the multiple second video segments, and the person relationship information corresponding to the hit video segment is consistent with the first information; the person relationship information is used to represent the relationship between the persons included in the visual media and the central person, and the central person refers to the person in the visual media stored in the electronic device who is in the social center; and displaying a first search result corresponding to the first video.

[0020] For the visual media search method provided in the second aspect, after the user enters the search text, the electronic device can identify the information of the search text, including the first information representing the person relationship in the search text. Then, the electronic device can search for the video corresponding to the video segment whose person relationship is consistent with the first information among the video segments included in each video to obtain the first video. On the one hand, this method can implement the search for videos, provide more and more comprehensive search results to the user, and improve the user experience. On the other hand, the user can search for matching videos through the person relationship, providing the user with more dimensions of search methods and improving the user experience. On the further hand, this method searches based on multiple video segments of the video. The video is dynamic, and there are often many scene changes in a video. Different video segments can represent different scenes. In this way, not only can multiple image semantics and multiple person relationships in the video be recognized, and the information obtained from the video is more comprehensive and accurate. Moreover, for each video segment, the disturbance of irrelevant information can be reduced, improving the accuracy of the search.

[0021] In a possible implementation, the method further includes: searching for a first image among multiple images according to the first information; the person relationship information corresponding to the first image is consistent with the first information; displaying a second search result corresponding to the first image.

[0022] That is to say, the search results include not only videos but also images. In this way, more and more comprehensive search results are provided to the user, improving the user experience.

[0023] In a possible implementation, the method further includes: performing natural language understanding processing on the search text to identify a first semantic entity and category information in the search text; performing text semantic understanding on the search text to obtain a text semantic vector; searching for a first video among multiple videos according to the first information, including: performing multi-way recall in the index library according to the first information, the first semantic entity, the category information, and the text semantic vector to obtain a recall result, and the recall result includes the first video; the index library includes index of attribute information, person relationship information, and visual semantic vectors of multiple video segments and multiple images, and the attribute information includes at least one of acquisition time of the visual media, acquisition location of the visual media, classification label of the visual media, and semantic entity of the visual media.

[0024] In this implementation, matching is performed with the search text from multiple dimensions, that is, searching comprehensively from multiple dimensions, greatly improving the matching degree between the search results and the search text, and improving the user experience. Moreover, sorting the multi-way recall results enables the user to preferentially see visual media with a high matching degree with the search text, improving the user experience.

[0025] In a possible implementation, the method further includes: receiving a third operation for the user to add and / or modify the target visual media; the target visual media includes a target video and a target image; in response to the third operation, performing segmentation processing on the target video to obtain multiple third video segments; extracting a frame image from each third video segment as a representative frame image; respectively obtaining the person relationship information and visual semantic vectors of each target image and each representative frame image; respectively obtaining the attribute information of each target image and the attribute information of each target video; based on the person relationship information and visual semantic vectors of each target image and each representative frame image, and the attribute information of each target image and each target video, constructing indexes of the attribute information, person relationship information, and visual semantic vectors of each target image and each third video segment in the index library.

[0026] In this implementation, during processing such as person relationship analysis and semantic understanding, it is based on the representative frames extracted from each video segment. In this way, it is not necessary to process each frame image in the video segment separately, which can not only reduce noise information, but also reduce the amount of calculation, storage, etc., thereby reducing resource overhead.

[0027] In a possible implementation, the person relationship information and visual semantic vectors of each target image and each representative frame image are obtained respectively, including: performing person relationship analysis on each target image and each representative frame image respectively to obtain the person relationship information corresponding to each target image and the person relationship information corresponding to each representative frame image; performing image semantic understanding on each target image and each representative frame image respectively to obtain the visual semantic vector of each target image and the visual semantic vector of each representative frame image.

[0028] In a possible implementation, based on the person relationship information and visual semantic vectors of each target image and each representative frame image, and the attribute information of each target image and each target video, indexes of the attribute information, person relationship information, and visual semantic vectors of each target image and each third video segment are constructed in the index library, including: using the person relationship information corresponding to the first representative frame image as the person relationship information corresponding to the fourth video segment, where the first representative frame image is any one of the representative frame images, and the fourth video segment is the video segment to which the first representative frame image belongs; using the visual semantic vector of the first representative frame image as the visual semantic vector of the fourth video segment; based on the person relationship information and visual semantic vectors of each target image and each third video segment, and the attribute information of each target image and each target video, constructing indexes of the attribute information, person relationship information, and visual semantic vectors of each target image and each third video segment in the index library.

[0029] In a possible implementation, performing image semantic understanding on each target image and each representative frame image respectively to obtain the visual semantic vector of each target image and the visual semantic vector of each representative frame image, including: searching for person relationship information from the annotation information of the target image and the target video by the user to obtain the annotation relationship information; identifying the social circle and the type of the social circle according to the face image set; the face image set includes multiple face images, a face image is an image containing a face, the multiple face images include the target face image and the face representative frame, the target face image is the image containing a face in the target image, the face representative frame is the image containing a face in the representative frame, the social circle is a set of people having the same social relationship with the central person, and the type of the social circle represents the type of the social relationship between the people in the social circle and the central person; determining the person relationship information corresponding to each target image and the person relationship information corresponding to each representative frame image according to the annotation relationship information, the social circle, and the type of the social circle.

[0030] In a possible implementation, based on a set of face images, a social circle and the type of the social circle are identified, including: performing face clustering on multiple face images to generate a face clustering result; the face clustering result representing the corresponding relationships among multiple face images, multiple faces, and multiple persons; determining a heterogeneous graph of person images according to the face clustering result; the heterogeneous graph of person images representing the corresponding relationships between multiple face images and multiple persons; determining the intimacy between every two persons among multiple persons according to the heterogeneous graph of person images, the intimacy representing the degree of association between persons; identifying the social circle according to the intimacy between every two persons and the heterogeneous graph of person images; and determining the type of the social circle according to the heterogeneous graph of person images.

[0031] In this implementation, on the one hand, based on the face clustering result, the social circle and the type of the social circle can be identified without relying on user annotation, with high intelligence and improved user experience. On the other hand, this process can accurately identify the central person without obtaining information such as the user's face unlocking, and then identify the social circle and the type of the social circle, without affecting the security of the user's private information and further improving the user experience. On the third hand, the above process can be completed offline based on the images in the gallery database without relying on the network. Compared with identifying the social circle and the type of the social circle through social behaviors on the social network, this method has stronger applicability, and the recognition result is more targeted for the user of the electronic device, thus further improving the user experience. On the fourth hand, the above process first identifies the social circle, then determines the type of the social circle based on the recognition result of the social circle, and corrects the relationship types of the persons in the social circle. Compared with separately classifying the relationships between persons to obtain different social circles and the types of social circles, the method provided in this application has higher accuracy in identifying the social circle and the type of the social circle and better user experience.

[0032] In a possible implementation, the heterogeneous graph of person images includes multiple image nodes and multiple person nodes, the multiple image nodes corresponding one-to-one to multiple face images, and the multiple person nodes corresponding one-to-one to multiple persons; identifying the social circle according to the intimacy between every two persons and the heterogeneous graph of person images, including: removing the multiple image nodes in the heterogeneous graph of person images, and connecting the person nodes corresponding to the two persons with non-zero intimacy by an edge according to the intimacy between every two persons to obtain a person connection graph; identifying the central person according to the heterogeneous graph of person images; removing the person node corresponding to the central person in the person connection graph to obtain a de-centered person connection graph; and performing community discovery on the de-centered person connection graph to obtain the social circle.

[0033] In this implementation, the heterogeneous graph of person images is simplified to a person connection graph, and there is no need to find the relationships between persons according to the images in the gallery application, simplifying the process of obtaining the person connection graph and improving the algorithm operation efficiency.

[0034] In a possible implementation, according to the heterogeneous graph of person images, the type of social circle is determined, including: inputting the heterogeneous graph of person images into a graph neural network model to extract the person feature vectors of each person among multiple persons and the image feature vectors of each face image; inputting the person feature vectors and the image feature vectors into a multi-layer perceptron model to predict the type of social relationship between each person among multiple persons and the central person; determining the number of persons corresponding to each type of social relationship in the social circle; and determining the type of social relationship with the largest corresponding number of persons as the type of the social circle.

[0035] In this implementation, based on the pre-trained graph neural network model and multi-layer perceptron model, it is possible to simply, quickly, and accurately predict the type of social relationship between a person and the central person, and then determine the type of social circle according to the type of social relationship.

[0036] In a possible implementation, among the visual media stored in the electronic device, the time distribution divergence and location distribution divergence of the visual media containing the central person satisfy preset conditions. The time distribution divergence is the distribution divergence of the acquisition time of the visual media, and the location distribution divergence is the distribution divergence of the acquisition location of the visual media.

[0037] The central person is at the center of the social interaction. For any person, the more uniform the location distribution and the more uniform the time distribution, the greater the possibility that this person is the central person. In this implementation, the central person determined by whether the time distribution divergence and location distribution divergence satisfy the preset conditions is more accurate.

[0038] In a third aspect, the present application provides a device, which is included in an electronic device and has the function of implementing the behavior of the electronic device in the first aspect and the possible implementation manners of the first aspect above, or has the function of implementing the behavior of the electronic device in the second aspect and the possible implementation manners of the second aspect above. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions. For example, a receiving module or unit, a processing module or unit, etc.

[0039] In a fourth aspect, the present application provides an electronic device, which includes: a processor, a memory, and an interface; the processor, the memory, and the interface cooperate with each other to enable the electronic device to execute any one of the methods in the technical solutions of the first aspect and the second aspect.

[0040] In a fifth aspect, the present application provides a chip, including a processor. The processor is used to read and execute a computer program stored in a memory to execute the method in the first aspect and any of its possible implementation manners, or execute the method in the second aspect and any of its possible implementation manners.

[0041] Optionally, the chip further includes a memory, and the memory is connected to the processor through a circuit or a wire.

[0042] Further optionally, the chip further includes a communication interface.

[0043] In a sixth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the processor is caused to execute any one of the methods in the technical solutions of the first aspect and the second aspect.

[0044] In a seventh aspect, the present application provides a computer program product, which includes: computer program code. When the computer program code runs on an electronic device, the electronic device is caused to execute any one of the methods in the technical solutions of the first aspect and the second aspect. Description of the Drawings

[0045] Figure 1 is a schematic diagram of an application scenario of a visual media search method provided by an embodiment of the present application;

[0046] Figure 2 is a schematic diagram of another application scenario of a visual media search method provided by an embodiment of the present application;

[0047] Figure 3 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application;

[0048] Figure 4 is a software structure block diagram of an electronic device provided by an embodiment of the present application;

[0049] Figure 5 is a schematic diagram of the interface change of a visual media search method provided by an embodiment of the present application;

[0050] Figure 6 is a schematic diagram of the interface change of another visual media search method provided by an embodiment of the present application;

[0051] Figure 7 is a schematic diagram of a search entry provided by an embodiment of the present application;

[0052] Figure 8 is a schematic diagram of the principle of a visual media search method provided by an embodiment of the present application;

[0053] Figure 9 is a process interaction diagram of a visual media search method provided by an embodiment of the present application;

[0054] Figure 10 is a process interaction diagram of another visual media search method provided by an embodiment of the present application;

[0055] Figure 11 It is a schematic flowchart of a video segmentation process provided by an embodiment of the present application;

[0056] Figure 12 It is a schematic diagram of the principle of a video segmentation process provided by an embodiment of the present application;

[0057] Figure 13 It is a schematic flowchart of a person relationship analysis provided by an embodiment of the present application;

[0058] Figure 14 It is a schematic logical diagram of a person relationship analysis provided by an embodiment of the present application;

[0059] Figure 15 It is a person-image heterogeneous graph provided by an embodiment of the present application;

[0060] Figure 16 It is a person connection diagram provided by an embodiment of the present application;

[0061] Figure 17 It is a decentralized person connection diagram provided by an embodiment of the present application;

[0062] Figure 18 It is a schematic diagram of a scenario for removing edge person nodes provided by an embodiment of the present application;

[0063] Figure 19 It is a person-social circle graph provided by an embodiment of the present application;

[0064] Figure 20 It is a schematic diagram of the type recognition results of a social circle and a social circle provided by an embodiment of the present application;

[0065] Figure 21 It is a co-occurrence time distribution graph of person ab provided by an embodiment of the present application;

[0066] Figure 22 It is a comparison schematic diagram of the co-occurrence time distribution of person ab and the uniform time distribution provided by an embodiment of the present application;

[0067] Figure 23 It is a co-occurrence location distribution graph provided by an embodiment of the present application;

[0068] Figure 24 It is a comparison schematic diagram of the co-occurrence location distribution of person ab and the uniform location distribution provided by an embodiment of the present application;

[0069] Figure 25 It is a schematic diagram of the interface for recommended image collections provided by an embodiment of the present application. Detailed implementation manners

[0070] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; "and / or" herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality" means two or more than two.

[0071] Hereinafter, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include one or more of such features.

[0072] Referring to "one embodiment" or "some embodiments" described in the specification of the present application means that specific features, structures, or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" appearing in different places in the specification of the present application do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0073] To better understand the embodiments of the present application, the following explains the terms or concepts that may be involved in the embodiments.

[0074] 1. Visual media

[0075] Visual media refers to the media that transmits information through vision. In the embodiments of the present application, visual media may include images, videos, etc.

[0076] 2. Semantic entity

[0077] The semantic entity can also be called an entity. Semantic entities include but are not limited to: entities with specific meanings such as time, person names (PER), and place names (LOC). Optionally, semantic entities can be identified by named entity recognition technology (NER).

[0078] 3. Text semantic vector

[0079] A text semantic vector refers to a vector that can represent the semantics of an entire text.

[0080] 4. Visual semantic vector

[0081] A visual semantic vector refers to a vector that can represent the visual semantics of visual media. In the embodiments of the present application, a visual semantic vector refers to a vector that can represent the semantics of an image or the semantics of a video segment.

[0082] 5. Vector similarity

[0083] Vector similarity is used to describe the similarity degree between two vectors. Generally, vector similarity can be calculated through the cosine similarity calculation formula. Of course, it can also be calculated through other methods.

[0084] 6. Heterogeneous graph

[0085] A heterogeneous graph refers to a graph in which the sum of the node types and the edge types is greater than 2. That is, a graph with node type + edge type > 2. It can be understood that a heterogeneous graph belongs to graph data, and its manifestation form can be an image, a list, a vector, a class, etc. The embodiments of the present application do not make any limitations in this regard. Specifically, relevant examples in subsequent embodiments can be referred to.

[0086] 7. Homogeneous graph

[0087] A homogeneous graph refers to a graph in which both the node type and the edge type are only one type. Similarly, it can be understood that a homogeneous graph also belongs to graph data, and its manifestation form can be an image, a list, a vector, a class, etc. The embodiments of the present application do not make any limitations in this regard.

[0088] 8. Pointwise mutual information (PMI)

[0089] Mutual information is used to measure the correlation between two things. The greater the mutual information, the greater the correlation between the two things. The smaller the mutual information, the smaller the correlation between the two things.

[0090] 9. Acquisition time

[0091] The acquisition time information of an image refers to the information related to the acquisition time of the image. Optionally, the acquisition time information may include the acquisition moment, the week information corresponding to the acquisition moment (also referred to as week information), etc. Among them, the acquisition moment may include year information, month information, day information, hour information, minute information, second information, etc. It can be understood that depending on different timing precisions, the types of information included in the acquisition moment can be different.

[0092] For ease of understanding, in the embodiments of the present application, the year information, month information, and day information are collectively referred to as date information, and the hour information, minute information, and second information are collectively referred to as timing information. For example, if the acquisition time of Image A is 12 seconds, 18 minutes, 17 hours, April 12th, 2023, then the year information included in this acquisition time is 2023, the month information is April, the day information is the 12th, the hour information is 17 hours, the minute information is 18 minutes, and the second information is 12 seconds. Among them, the date information is April 12th, 2023, and the timing information is 17 hours, 18 minutes, and 12 seconds. Similarly, it can be understood that with different timing accuracies, the types of information included in the date information and timing information can be different.

[0093] 10. Temporal distribution of images

[0094] The embodiments of the present application involve the temporal distribution of images. The temporal distribution of images characterizes the distribution of the acquisition times of images, that is, according to the acquisition time, the distribution of the number of images in different time intervals is determined. For example, according to the timing information in the acquisition time, the number of images in the time interval TA is pic1, and the number of images in the time interval TA2 is pic2.

[0095] 11. Location distribution

[0096] The embodiments of the present application involve the location distribution of images. The location distribution of images characterizes the distribution of the acquisition locations of images, that is, according to the acquisition location, the distribution of the number of images at different locations is determined. For example, the number of images with the acquisition location determined to be L1 is pic3, and the number of images with the acquisition location determined to be L2 is pic4.

[0097] 12. Divergence

[0098] Divergence is a data distribution characteristic and is an index commonly used in statistics to describe the group structure, which is used to represent the degree of deviation of data from the mean. Divergence can characterize the uniformity of data distribution.

[0099] The temporal distribution divergence refers to the degree of deviation of the temporal distribution from the mean of the temporal distribution (i.e., the uniform temporal distribution). The temporal distribution divergence can reflect the uniformity of the temporal distribution. The smaller the temporal distribution divergence, the more uniform the temporal distribution; the larger the temporal distribution divergence, the more uneven the temporal distribution, that is, the more concentrated.

[0100] The location distribution divergence refers to the degree of deviation of the location distribution from the mean of the location distribution (i.e., the uniform location distribution). The location distribution divergence can reflect the uniformity of the location distribution. The smaller the location distribution divergence, the more uniform the location distribution; the larger the location distribution divergence, the more uneven the location distribution, that is, the more concentrated.

[0101] 13. KL divergence

[0102] KL is the abbreviation of kullback - leibler. KL divergence is used to measure the difference between two distributions. The larger the KL divergence, the greater the difference between the two distributions. The smaller the KL divergence, the smaller the difference between the two distributions. The value of KL divergence can be in the range of [0, 1]. When the two distributions are the same, the KL divergence is 0.

[0103] 14. Co - occurrence relationship

[0104] Co - occurrence means occurring together. In the embodiments of the present application, the co - occurrence relationship refers to the co - occurrence relationship between people in an image. That is to say, when people appear together in an image, it is called the co - occurrence relationship between people.

[0105] Before elaborating on the visual media search method provided by the embodiments of the present application in detail, first, the application scenario of this method and the structure of the applicable electronic device will be described.

[0106] Optionally, the visual media search method provided by the embodiments of the present application can be applied to the gallery application program (APP) of an electronic device, and is used to search for images and videos that match the search content input by the user. The following will be described in conjunction with the interface diagram.

[0107] Exemplarily, Figure 1 FIG. (a) shows a schematic diagram of an application scenario of a visual media search method provided by an embodiment of the present application. Taking the electronic device as a mobile phone as an example, as Figure 1 shown in FIG. (a), the main interface of the mobile phone includes the icon 101 of the gallery APP. The user clicks on this icon 101, and in response to the user's click, the mobile phone opens the gallery APP and displays Figure 1 the album interface shown in FIG. (b). The album interface 102 includes a search box 103. The user can enter the search content in the search box 103. For example, the user can enter "Beijing" in the search box 103. In response to the user's operation, the mobile phone can search for images with the shooting location of "Beijing" in the gallery and display the search results, as Figure 1 shown in FIG. (c).

[0108] In the related art, the search function can only be applied to the search for images and is not applicable to the search for videos. Specifically, as Figure 1 shown in FIG. (c), the search results do not include videos. Moreover, in the related art, the search is mainly based on the shooting location, shooting time, the names of people or other tags labeled by the user for the image, etc., and the content without labeled tags cannot be searched, so the accuracy of the search results is not high and the user experience is poor. In addition, in the related art, it is impossible to search for images or videos based on the relationship between people, and the search experience brought to users needs to be improved. For example, as Figure 2As shown in Figure (a) therein, when the user enters "having dinner with family members" in the search box 103 of the photo album interface 102, even if the picture library contains images that match this statement, the mobile phone cannot search for results, as Figure 2 shown in Figure (b) therein.

[0109] In view of this, the embodiment of the present application provides a visual media search method. On the one hand, this method can not only be applicable to images, but also to videos, providing users with more and more comprehensive search results and improving the user experience. On the other hand, this method performs search and matching based on semantic vectors. Therefore, users can perform fuzzy searches for statements. For example, users can enter search statements such as "a woman doing yoga", and the electronic device can perform a search based on this statement, and can more accurately search for visual media that meets the actual intentions of users, improving the search precision and the user experience. On the third hand, this method can identify the relationship information of people in visual media and establish a corresponding relationship between visual media and the relationship of people. In this way, when the user searches for a certain type of relationship of people, it can accurately display images and / or videos that have this type of relationship with the user (such as the owner of the device) to the user, providing users with more dimensional search methods and improving the user experience.

[0110] The visual media search method provided by the embodiment of the present application can be applied to electronic devices such as mobile phones, tablet computers, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc. that can install application programs (APPs). The embodiment of the present application does not impose any restrictions on the specific types of electronic devices.

[0111] Exemplarily, Figure 3This is a schematic structural diagram of an electronic device 100 provided by an embodiment of the present application. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0112] It can be understood that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0113] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0114] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching instructions and executing instructions.

[0115] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may hold instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses and reduces the waiting time of the processor 110, thus improving the efficiency of the system.

[0116] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0117] It can be understood that the interface connection relationships among the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0118] The charging management module 140 is configured to receive a charging input from a charger. The charger may be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 may receive the charging input from the wired charger through the USB interface 130. In some embodiments of wireless charging, the charging management module 140 may receive the wireless charging input through the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 may also supply power to the electronic device through the power management module 141.

[0119] The electronic device 100 may implement a shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.

[0120] The software system of the electronic device 100 may adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of this application, taking the Android system with a layered architecture as an example, the software structure of the electronic device 100 will be exemplarily described.

[0121] Figure 4 It is the software structure block diagram of the electronic device 100 in the embodiments of this application. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0122] The application layer may include a series of application packages. As Figure 3 shown, the application packages may include a gallery APP, a search module, a natural language understanding module, and computer vision (CV) services, etc.

[0123] The gallery APP is also referred to as a gallery service or a gallery service module, etc., and is used to provide services such as storage, management, display, and recommendation of visual media such as images and videos to users. Optionally, the gallery APP may have a database, hereinafter referred to as the gallery database. The gallery database is used to store relevant data of the gallery APP, such as images and their attribute information, videos and their attribute information, visual semantic vectors, face clustering results, person relationship information, and so on.

[0124] The search module is used to search for visual media according to the content input by the user. Optionally, the search module may have its own database for storing index data. Among them, the indexing method of the index data may be an inverted index, and hereinafter the database of the search module will be referred to as the inverted index library.

[0125] The natural language understanding module is used to understand the information contained in the text. In the embodiments of this application, the natural language understanding module may be used to identify semantic entities, person relationships, category information, etc. in the text. The category information refers to the information characterizing the type to which an image or video belongs. Exemplarily, the category information may be a person, a plant, an animal, a building, or a natural scenery, etc.

[0126] In one embodiment, as Figure 3 shown, the CV services may include a face analysis module, a person relationship analysis module, a video segmentation module, and a multimodal understanding module, etc.

[0127] The face analysis module is used to perform face analysis on visual media. Among them, face analysis includes but is not limited to face recognition, face clustering, gender recognition, age recognition, etc.

[0128] The person relationship analysis module is used to analyze the social relationships between the persons included in the visual media to obtain person relationship information.

[0129] The video segmentation module is used to perform segmentation processing on a video, divide the video into multiple video segments according to semantic content, and extract one frame image from each video segment as a representative frame (also referred to as a representative frame image). The representative frame is used to represent the semantics of the video segment.

[0130] The multi-modal understanding module, also known as the multi-modal semantic understanding module, multi-modal semantic understanding model, etc., is used to perform semantic understanding on one or more of various modal information such as text, image, video, audio, etc. In the embodiments of the present application, the multi-modal understanding module can be used to perform semantic understanding on an image to obtain a visual semantic vector, and can also be used to perform semantic understanding on text (such as search text) to obtain a text semantic vector.

[0131] Optionally, the CV service can have its database, hereinafter referred to as the CV database. The CV database is used to store the data required for the operation of each module in the CV service, such as storing images containing persons, face clustering results, person relationship analysis results, and so on.

[0132] Of course, in addition to Figure 4 each module or APP shown in, the application layer may also include applications such as a camera, a call, a map, a Bluetooth, a short message, etc. (not shown in the figure), and the embodiments of the present application do not make any limitation thereto.

[0133] The application framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer includes some predefined functions.

[0134] The application framework layer may also include a window manager, a content provider, a view system, a telephone manager, a resource manager, a notification manager, etc.

[0135] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.

[0136] The content provider is used to store and obtain data, and make these data accessible to applications. The data may include videos, images, audios, incoming and outgoing calls, browsing history and bookmarks, phone books, etc.

[0137] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a text message notification icon can include a view for displaying text and a view for displaying pictures.

[0138] The phone manager is used to provide the communication functions of the electronic device 100. For example, the management of call status (including answering, hanging up, etc.).

[0139] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, and so on.

[0140] The notification manager enables applications to display notification information in the status bar. It can be used to convey notification-type messages, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that the download is complete, message reminder, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as the notification of a background-running application, or a notification that appears in the form of a dialogue window on the screen. For example, it prompts text information in the status bar, emits a prompt tone, the electronic device vibrates, the indicator light flashes, etc.

[0141] Optionally, each application package in the application layer can implement the related functions of the gallery APP by calling the algorithms or modules in the application framework layer. For example, when the face analysis module performs face clustering, it can call the face clustering algorithm module (not shown in the figure) in the application framework layer to cluster the faces in the image and obtain the face clustering result. Another example is that when the person relationship analysis module performs person relationship analysis, it can call the social circle recognition algorithm module (not shown in the figure) in the application framework layer, and based on the face clustering result, identify the central person and discover the social circle of the central person. The detailed implementation of the above process will be further elaborated in the subsequent embodiments.

[0142] Android runtime includes a core library and a virtual machine. Android runtime is responsible for the scheduling and management of the Android system.

[0143] The core library contains two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core library of Android.

[0144] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files in the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.

[0145] The system library may include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing library (e.g., OpenGL ES), 2D graphics engine (e.g., SGL), etc.

[0146] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.

[0147] The media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0148] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.

[0149] The 2D graphics engine is a drawing engine for 2D drawing.

[0150] The kernel layer is the layer between hardware and software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.

[0151] For ease of understanding, the following embodiments of the present application will take an electronic device with the Figure 3 and Figure 4 shown structure as an example, and in combination with the accompanying drawings and application scenarios, specifically elaborate on the visual media search method provided by the embodiments of the present application.

[0152] First, the interface effect of the visual media search method provided by the embodiments of the present application will be described.

[0153] Exemplarily, Figure 5 is a schematic diagram of the interface change of an example of the visual media search method provided by the embodiments of the present application. As shown in Figure 5 in figure (a), the album interface 102 of the gallery APP is displayed on the mobile phone screen. The album interface 102 includes a search box 103. The user enters the search text "boys standing outdoors" in the search box 103. The mobile phone responds to the user's operation, searches the gallery database for images and videos matching "boys standing outdoors", and displays the search results, as shown in Figure 5As shown in the search result interface 501 of FIG. (b). The search result interface 501 may include a first display area 502, a second display area 503, and a third display area 504. In the first display area 502, some search results (e.g., 3) are shown, including images and videos. In the second display area 503, a video that best matches the search text in the search results is shown, and the number of videos included in the search results, "4", is shown. In the third display area 504, an image that best matches the search text in the search results is shown, and the number of images included in the search results, "6", is shown. It can be understood that in the embodiments of the present application, the search results may be shown in the form of thumbnails, that is, the thumbnails of the images are shown, or the thumbnails of a certain frame image of the video are shown as the thumbnails of the video.

[0154] It can be seen that in the embodiments of the present application, when the search content is a sentence, visual media that semantically matches the sentence can also be obtained, that is, the method can perform fuzzy search based on the sentence. Moreover, in addition to images, the search results also include videos.

[0155] Specifically, for videos, the method pre-processes each video by segmenting it, and divides the video into multiple video segments according to semantic content. During the search, the search text is matched with the semantics of each video segment. If there is at least one video segment in a certain video that semantically matches the search text, then that video is shown as a search result. In subsequent embodiments, the video segment in the video that matches the search text is called a hit video segment.

[0156] In one embodiment, the cover of the video in the search results may show relevant information of the video. For example, as Figure 5 shown in FIG. (b), the cover of the video 505 may show the total duration of the video, 8 minutes and 32 seconds (08:32).

[0157] In one embodiment, the cover of the video in the search results may also show the start playing time (abbreviated as start time) of the hit video segment. As Figure 5 shown in FIG. (b), taking the start time of the hit video segment in the video 505 as 2 minutes and 18 seconds (02:18) and the end time as 3 minutes and 18 seconds (03:18) as an example, the cover of the video 505 may show the start time 2 minutes and 18 seconds (02:18).

[0158] In some other embodiments, the cover of the video may also show the middle playing time (abbreviated as middle time) of the hit video segment in the video. Continuing with the above example, the cover of the video may also show the middle time 2 minutes and 48 seconds (02:48) from 2 minutes and 18 seconds (02:18) to 3 minutes and 18 seconds (03:18).

[0159] In still other embodiments, after the video is divided into video segments, a frame image can be selected from each video segment as a representative frame, and the cover of the video can also display the playback time corresponding to the representative frame of the video segment hit in the video. For example, if the playback time corresponding to the representative frame of the hit video segment is 2 minutes and 0 seconds (02:00), the cover of the video can display this time.

[0160] In some embodiments, the cover image of the video in the search results can be the first frame image of the video. In other embodiments, the cover image of the video in the search results can also be the first frame image, the last frame image, the middle frame image, or the representative frame image of the hit video segment, etc., and the embodiments of the present application do not make any limitations in this regard.

[0161] In addition, if multiple video segments in the video are hit, one of the multiple hit video segments can be selected, and the information of this video segment can be displayed on the video cover in the above manner. In one embodiment, the multiple hit video segments can be sorted according to the matching degree with the search text, and the information of the hit video segment with the highest matching degree with the search text can be displayed on the video cover in the above manner.

[0162] In some embodiments, when displaying the search results, the videos and images in the search results can also be sorted according to the matching degree with the search text. Optionally, the sorting principle can be: the higher the matching degree with the search text, the higher the sorting during display.

[0163] In summary, the method provided by the embodiments of the present application can have various display methods for the search results, and the embodiments of the present application do not make any limitations in this regard.

[0164] Next, taking the search text including a relationship between people as an example, the effect interface of the method provided by the embodiments of the present application will be described.

[0165] Exemplarily, Figure 6 is a schematic diagram of the interface change of another visual media search method provided by the embodiments of the present application. As shown in Figure (a) of Figure 6 , the user enters the search text "traveling with family last year" in the search box 103 on the album interface 102 of the mobile phone. Among them, "family" belongs to the relationship information between people. The mobile phone responds to the user's operation, searches the gallery database for images and videos matching "traveling with family last year", and displays the search results in the search result interface 601, as shown in Figure (b) of Figure 6 . It can be seen that in the embodiments of the present application, when the search text contains a relationship between people, images and / or videos that conform to this relationship between people can also be searched, meeting the user's expectations and improving the user experience.

[0166] In this embodiment, the display method of the images and videos in the search results can be the same as that in the aboveFigure 5 This is the same as the illustrated embodiment and will not be elaborated further.

[0167] As Figure 6 shown in figure (b) of Figure 6 , taking the search results including video 602 as an example, the user can click on video 602, and in response to the user's operation, the mobile phone jumps to the video playback interface 603, as Figure 6 shown in figure (c) of

[0168] In one embodiment, the mobile phone can start playing the video from the start time of the hit video segment, as Figure 6 shown in figure (d) of

[0169] In another embodiment, it is also possible to start playing the video from the playback time corresponding to the representative frame of the hit video segment.

[0170] In yet another embodiment, it is also possible to start playing the video from the start time (00:00) of video 602.

[0171] In yet another embodiment, when video 602 includes multiple hit video segments, the hit video segments can be played sequentially in the chronological order of the video segments in video 602, or can be played sequentially according to the degree of matching with the search text. The embodiments of the present application do not make any limitation in this regard.

[0172] It can be understood that Figure 6 in

[0173] the effect of the visual media search method is illustrated by taking the search text including the "family member" person relationship information as an example. It can be understood that the search text can also include other person relationships, such as "colleague", "friend", "classmate", "son", "daughter", "best friend", etc. Figure 7 In addition, the above embodiments are all illustrated by taking the search in the album interface 102 in the photo gallery as an example. In practical applications, the search can also be implemented in other interfaces of the photo gallery APP, or in other application or service interfaces. For example, referring to Figure 7 figure (a) of Figure 7In the figure (c), you can also input content in the global search box 706 on the negative first screen interface 705 of the mobile phone to achieve search. The embodiments of the present application do not make any limitations on the search interface and search entry.

[0174] Next, in combination with the interaction diagram, corresponding to the above interface changes, the implementation process of the visual media search method provided by the embodiments of the present application will be described.

[0175] This method may include two stages: index construction and search. Among them, the index construction stage belongs to the offline stage, and the search stage belongs to the online stage.

[0176] Exemplarily, Figure 8 is a schematic diagram of the principle of a visual media search method provided by an embodiment of the present application. As Figure 8 shown, in the index construction stage, for the videos in the gallery database, after frame splitting processing, multiple frame images are obtained; the multiple frame images are segmented to obtain multiple video segments; representative frames are extracted from the video segments, and the content of the video segments is represented by the representative frames; for the images in the gallery database, no frame splitting and segmentation processing are required. Therefore, the images in the gallery database and the representative frames obtained after segmenting the videos are analyzed for the relationship between characters to obtain character relationship information, and semantic understanding is performed on these images and representative frames (for example, input into an image semantic understanding model) to obtain corresponding visual semantic vectors; an inverted index is constructed based on the obtained visual semantic vectors and character relationship information, and the inverted index is stored in the inverted index library.

[0177] In the search stage, natural language understanding is performed on the search text input by the user to identify semantic subjects, character relationships, category information, etc. contained in the search text; in addition, semantic understanding is performed on the search text (for example, input into a text semantic understanding model) to obtain a corresponding text semantic vector; multi-way recall is performed from the inverted index library based on the semantic subject, character relationship, category information, and text semantic vector; then, the multi-way recall results are filtered and sorted to obtain search results; the search results are displayed on the search interface.

[0178] Optionally, when training the above image semantic understanding model and text semantic understanding model, they can be trained in the direction of contrast learning to make the semantic understanding of the two more consistent and improve the accuracy of visual media search.

[0179] The specific implementation processes of the two stages will be described separately below.

[0180] 1. Index construction stage

[0181] Figure 9 is a process interaction diagram of a visual media search method provided by an embodiment of the present application. As Figure 9 shown, the index construction stage of this method may include:

[0182] S101. In response to the user adding and / or modifying visual media, the gallery APP stores the incremented visual media added and / or modified by the user and its attribute information in the gallery database.

[0183] Optionally, the user can add visual media by means such as shooting, downloading, taking screenshots, etc. In addition, the user can also modify the existing visual media to obtain new visual media. Such modifications include but are not limited to: beautifying, editing, custom adding names of people, adding watermarks, etc.

[0184] The attribute information of the visual media can include but is not limited to: identification information, acquisition location, acquisition time, names of people, tags (also known as classification tags). Optionally, the identification information of the visual media can be represented by an identity document number (ID) (such as a hash ID), or can also be represented by a name, path, etc. This application does not make any limitations in this regard. In the following embodiments, taking the identification information of the visual media being represented by an ID as an example for illustration.

[0185] The acquisition location refers to the location where the visual media is acquired. The acquisition time refers to the time when the visual media is acquired. For a captured video or image, the acquisition location can be the shooting location, and the acquisition time can be the shooting time. In one embodiment, for a captured video or image, the attribute information can also include the camera information used for shooting (such as whether it is shot with the front camera). For a screenshot, the acquisition location can be the screenshot location, and the acquisition time can be the screenshot time. For a downloaded video or image, the acquisition location can be the download location, and the acquisition time can be the download time.

[0186] The name of a person refers to the name added by the user to the person in the image or video. A tag refers to information characterizing the type or attribute of an image or video. Exemplarily, tags can be people, plants, animals, buildings, or natural scenery, etc.

[0187] For the target visual media added by the user and its attribute information, the gallery APP stores the added visual media and its attribute information in the gallery database. For the visual media and its attribute information modified by the user, the gallery APP synchronously modifies the visual media and its attribute information stored in the gallery database.

[0188] In some other embodiments, the gallery APP can store the added visual media and its attribute information in the cloud to relieve the storage pressure on the local electronic device.

[0189] For ease of description, in the following embodiments, the visual media added and / or modified by the user are referred to as incremental visual media or target visual media, the videos in the visual media added and / or modified by the user are referred to as incremental videos or target videos, and the images in the visual media added and / or modified by the user are referred to as incremental images or target images.

[0190] S102. The gallery APP synchronizes the incremental visual media and their attributes to the CV database of the CV service.

[0191] When the user adds and / or modifies visual media, the gallery APP can synchronize the incremental visual media and their attributes to the CV database, which facilitates subsequent modules of the CV service to obtain data as needed.

[0192] S103. When the electronic device is in a charging state and the screen is in a turned-off state, the gallery APP sends a segmentation request to the video segmentation module of the CV service.

[0193] The segmentation request is used to request the segmentation of the incremental video and obtain the representative frames of the video segments. Optionally, the incremental video or the ID of the incremental video can be carried in the segmentation request.

[0194] Optionally, the gallery APP can subscribe to the charger plugging and unplugging state from a relevant module in the electronic device, and when the charger plugging and unplugging state changes, this module sends a notification to the gallery APP. In this way, the gallery APP can know in real time whether the charger is in an inserted state, that is, know whether the electronic device is in a charging state. Similarly, the gallery APP can also subscribe to the screen state from a relevant module in the electronic device, and when the screen state changes, this module sends a notification to the gallery APP. In this way, the gallery APP can know in real time whether the screen is in a turned-off state.

[0195] In another embodiment, the gallery APP can also send a segmentation request to the video segmentation module when the electronic device is in a charging state, the screen is in a turned-off state, and the current time is within a preset time period. Optionally, the preset time period can be, for example, from 00:00 to 07:00 every day. In this way, it is possible to reduce the disturbance to the user and improve the user experience.

[0196] S104. In response to the segmentation request, the video segmentation module segments the incremental video to obtain multiple video segments and their attribute information, and extracts representative frames from each video segment.

[0197] Specifically, the video segmentation module can call a preset video segmentation algorithm to segment each incremental video. The video segmentation by the video segmentation algorithm mainly includes three processes: the first process is to perform frame splitting on the incremental video to split it into multiple frame images; the second process is to segment the multiple frame images obtained by frame splitting based on the semantics of the video. Each video segment obtained by segmentation can represent a scene respectively, and there is a certain correlation between the frame images within each video segment; the third process is to extract one frame image from each video segment as a representative frame.

[0198] The representative frame is used to represent the semantics of the video segment. In one embodiment, the first frame image, the last frame image, or the middle frame image of the video segment can be used as the representative frame. In another embodiment, the representative frame can also be determined by calculating parameters such as the jitter degree, clarity, label, and pixel transformation degree of the frame images in the video segment. The embodiments of the present application do not make specific limitations on the method for determining the representative frame.

[0199] Optionally, the attribute information of the video segment can include video segment identification information, start time, end time, label union, etc. The identification information of the video segment can be, for example, the video segment ID. The label union refers to the union of the labels (referred to as frame labels) of the frame images in the video segment.

[0200] The specific implementation manner of step S104 will be further described in subsequent embodiments.

[0201] S105. The video segmentation module returns the video segment, its attribute information, and the representative frame to the gallery APP.

[0202] S106. The gallery APP stores the video segment, its attribute information, and the representative frame in the gallery database.

[0203] S107. The gallery APP sends a face analysis request to the face analysis module in the CV service.

[0204] The face analysis request is used to request face analysis of the incremental image containing a face (hereinafter referred to as the incremental face image or the target face image) and the representative frame containing a face (hereinafter referred to as the face representative frame). Face analysis can include face recognition, face clustering, etc.

[0205] Optionally, the incremental face image and the face representative frame are carried in the face analysis request. Or, the identification information of the incremental face image and the identification information of the face representative frame can be carried in the face analysis request.

[0206] S108. In response to the face analysis request, the face analysis module performs face analysis on the incremental face image and the face representative frame to obtain a face analysis result, and the face analysis result includes a face clustering result.

[0207] Specifically, the face analysis module can perform face recognition on the incremental face images and face representative frames through a face recognition algorithm. Optionally, the face information can be represented by a face ID.

[0208] After that, on the one hand, the face analysis module can analyze the gender, age, etc. of the person based on the recognized face data.

[0209] On the other hand, the face analysis module, based on a face clustering algorithm, clusters the face information in the recognized incremental face images and face representative frames together with the historical face information to obtain multiple classes, thereby updating the face clustering result obtained from the analysis in the historical time period.

[0210] Among them, the historical face refers to the face recognized in the image during the historical time period. It can be understood that the CV database can store historical face information. The face analysis module can perform face clustering based on the face recognition results of the incremental face images and face representative frames, combined with the historical face information in the CV database. Face clustering is to cluster faces with similar features into one class. In this way, at least one class is obtained, and each class can correspond to a person. Optionally, the face analysis module can identify each person. For example, the person can be numbered, and an ID corresponding to each person can be set. That is to say, each person can be represented by a person ID. The person ID is also called tag_id.

[0211] It should be noted that when the face analysis module performs face clustering, a situation where clustering cannot be performed may occur, that is, for a certain face, there is no other face with similar features. In this case, the face clustering algorithm can set the person ID corresponding to this face to a preset identifier. For example, it is set to -1. That is to say, a person ID of -1 means that there is no corresponding class. During the subsequent algorithm operation, the person with the person ID as the preset identifier can be excluded and not included in the operation scope, and only the persons with the person ID not being this preset identifier are considered. This can improve the accuracy of person relationship recognition and improve the user experience.

[0212] Exemplarily, Table 1 shows an example of the face clustering result provided by the embodiment of the present application.

[0213] Table 1

[0214] Serial number Image ID Face ID Person ID 1 42d8166356c2d1ff9875f602e97ea293 0 -1 2 42d8166356c2d1ff9875f602e97ea293 1 ser_246762133257613 3 b4a0d50b46f4f8ff86a25916c2cedecc 0 ser_112952351578583 4 2b869431a9182637d2a12c8444e352ec 0 -1 5 2586222e6dac13d5630edac4a8662208 0 ser_111346622209036 6 b320cc0aca977684063b4e3069f6fa08 0 -1 7 c65dbc7358974489502964155ad4cac0 0 ser_275636865923451 8 c65dbc7358974489502964155ad4cac0 1 ser_246762133257613 9 26fe24ebd66f90b540cc79637efdc46d 0 ser_136254524879543 10 26fe24ebd66f90b540cc79637efdc46d 1 ser_246762133257616 11 26fe24ebd66f90b540cc79637efdc46d 2 ser_112952351578583 12 26fe24ebd66f90b540cc79637efdc46d 3 ser_275636865923451

[0215] Taking the clustering results with serial numbers 9, 10, 11, and 12 in Table 1 as examples, it can be seen that there are 4 faces in the image with image ID "26fe24ebd66f90b540cc79637efdc46d", and the corresponding face IDs of the 4 faces are 0, 1, 2, and 3 respectively. Clustering the 4 faces, the classes to which the 4 faces belong are obtained, that is, the person IDs corresponding to the 4 face IDs are obtained, which are "ser_136254524879543", "ser_246762133257616", "ser_112952351578583", and "ser_275636865923451" respectively.

[0216] It should be noted that Table 1 is only an example of the face clustering results and does not impose any limitation on this application. In practical applications, the face clustering results can include more or less content than Table 1, or can be in other forms of representation.

[0217] S109. The face analysis module returns the face analysis result to the gallery APP.

[0218] S110. The face analysis module stores the incremental face image and the face representative frame in the CV database and updates the face image set in the CV data.

[0219] It can be understood that each time the face analysis module receives a face analysis request, it saves the incremental face image and the face representative frame requested for analysis in the face analysis request, as well as their attribute information, in the CV database. In this way, the CV database can store all the images containing faces and the representative frames containing faces in the gallery database, and these images or representative frames can be stored in the form of a set. This is convenient for subsequent processing such as analyzing the relationship between people based on the face image set. The images in the face image set are also called face images.

[0220] In addition, after each face analysis, the face analysis module also stores the face analysis results (including face clustering results, person age, person gender, etc.) in the CV database, which is also convenient for subsequent processing such as analyzing the relationship between people.

[0221] S111. The gallery APP stores the face analysis result in the gallery database.

[0222] S112. The gallery APP sends a person relationship analysis request to the person relationship analysis module in the CV service.

[0223] The request for analyzing the relationship between characters is used to request the analysis of the relationship information of the incremental face image and the face representative frame. The relationship information between characters is used to characterize the type of relationship between the characters included in the image and the central character. The relationship between characters can be, for example, family member, son, daughter, parent, colleague, friend, etc. The central character is the character in the center of the social relationship. The central character is generally considered to be the owner of the device or a person with a high degree of intimacy with the owner of the device.

[0224] Optionally, the request for analyzing the relationship between characters may carry the identification information of the incremental face image and the face representative frame.

[0225] S113. In response to the request for analyzing the relationship between characters, the module for analyzing the relationship between characters performs the analysis of the relationship between characters based on the face clustering result and the set of face images in the CV database, and obtains the relationship information of the incremental face image and the face representative frame.

[0226] The analysis of the relationship between characters is also called the analysis of social relationship, the discovery of the relationship between characters, etc.

[0227] Optionally, on the one hand, the module for analyzing the relationship between characters can obtain the relationship information of the characters actively added by the user. On the other hand, the module for analyzing the relationship between characters can call the social circle recognition algorithm, based on the face clustering result, identify the central character among the characters included in the set of face images, and identify the social circle formed by the characters in the set of face images and the type of the social circle based on the central character. The social circle of a character refers to the set of characters with social relationships. The type of the social circle is used to characterize the relationship between the characters in the social circle and the central character. Optionally, the type of the social circle can include family member, son, daughter, parent, colleague, friend, etc. That is to say, the type of the social circle can be used as the relationship information between characters.

[0228] S114. The module for analyzing the relationship between characters returns the relationship information of the incremental face image and the face representative frame to the Gallery APP.

[0229] It can be understood that each time the module for analyzing the relationship between characters performs the analysis of the relationship between characters based on the face clustering result and the set of face images, it will obtain the relationship information of all the images in the set of face images. The module for analyzing the relationship between characters can save this relationship information and return the relationship information corresponding to the incremental face image and the face representative frame to the Gallery APP.

[0230] S115. The Gallery APP stores the relationship information of the incremental face image and the face representative frame in the gallery database.

[0231] S116. The Gallery APP sends an image semantic understanding request to the multi-modal understanding module in the CV service.

[0232] The image semantic understanding request is used to request the semantic understanding of the incremental image and the representative frame.

[0233] Optionally, the incremental image and the representative frame may be carried in the image semantic understanding request, or the identification information of the incremental image and the representative frame may be carried.

[0234] S117. In response to the image semantic understanding request, the multimodal understanding module performs semantic understanding on the incremental image and the representative frame to obtain a visual semantic vector.

[0235] Specifically, the multimodal understanding module may include a pre-trained image semantic understanding model. By inputting each incremental image and representative frame into the image semantic understanding model respectively, the corresponding semantic vector can be obtained. In this embodiment, the semantic vector corresponding to the image is referred to as the visual semantic vector.

[0236] S118. The multimodal understanding module returns the visual semantic vectors of the incremental image and the representative frame to the Gallery APP.

[0237] S119. The Gallery APP stores the visual semantic vectors of the incremental image and the representative frame in the gallery database.

[0238] It can be understood that the visual semantic vector of each representative frame is used to represent the semantics of the video segment to which it belongs and has a unique correspondence with the video segment. Therefore, the visual semantic vector of the representative frame is also the visual semantic vector of the video segment.

[0239] S120. The Gallery APP sends the attribute information, the character relationship information, and the visual semantic vectors of each video segment in the incremental image and the incremental video to the search module.

[0240] S121. The search module constructs an inverted index corresponding to the incremental visual media according to the attribute information, the character relationship information, and the visual semantic vectors of each video segment in the incremental image and the incremental video, and stores the inverted index in the inverted index library.

[0241] It can be understood that when constructing the index, the search module can also construct a forward index. In this embodiment, constructing an inverted index can improve the subsequent search efficiency.

[0242] Specifically, the attribute information, the character relationship information, the visual semantic vector, etc. are collectively referred to as information. For the incremental image, the search module can directly establish the correspondence between the incremental image and the information of the incremental image.

[0243] For incremental videos, the search module can establish the correspondence between video segments and the information of video segments, and establish the correspondence between each video segment and the incremental video it belongs to, as well as the attribute information of the incremental video it belongs to. For the character relationship information and visual semantic vectors in the information of video segments, they can be represented by the information of the representative frame. That is to say, the character relationship information of the representative frame is used as the character relationship information of the video segment. The visual semantic vector of the representative frame is used as the visual semantic vector of the video segment.

[0244] In one embodiment, the relationship between the incremental video and its information is shown in Table 2 below. Optionally, the attribute information of the incremental video can be stored in a document, which can be called the parent document. The information of each video segment included in the incremental video can be stored in a document, which can be called the subdocument. An association relationship between the parent document and the subdocument is established.

[0245] Table 2

[0246]

[0247]

[0248] It should be noted that Table 2 is only an example for understanding the solution, and does not represent actual data, nor the presentation form of actual data.

[0249] It can be understood that every time the user adds or modifies visual media, the process shown in the above steps S101 to S121 is executed. In this way, the inverted index library can contain the indexes of all visual media saved in the gallery database. In addition, when the user deletes the data media or attribute information in the gallery database, the gallery APP deletes the visual media or attribute information, and notifies the CV service. The search module in the CV service can delete the corresponding indexes in the inverted index library. In this way, the update of the data in the inverted index library is realized. The process of data deletion will not be described in detail here. All in all, the index data in the inverted index library is consistent with the visual media in the gallery data. In this way, an accurate search function can be provided to the gallery APP.

[0250] In the embodiment of the present application, on the one hand, this method not only establishes indexes for incremental images, enabling users to search for matching images in the subsequent search stage, but also, for incremental videos, establishes indexes, enabling users to search for matching videos in the subsequent search stage, providing users with more and more comprehensive search results and improving the user experience.

[0251] In a second aspect, the method performs semantic understanding on representative frames in incremental images and incremental videos to obtain visual semantic vectors, and constructs an index based on the visual semantic vectors. In this way, in the subsequent search stage, users can perform fuzzy searches through search statements, and can more accurately search for images and / or videos that meet the actual intentions of the users, improving the search precision and the user experience.

[0252] In a third aspect, the method analyzes the relationships between characters in representative frames in incremental images and incremental videos to obtain character relationship information, and constructs an index based on the character relationship information. In this way, in the subsequent search stage, users can search for matching images or videos through character relationships, providing users with more dimensions of search methods and improving the user experience.

[0253] In a fourth aspect, the method can segment a video based on semantics and then perform related processing based on the video segments without the need for users to manually segment the video. Moreover, since a video is dynamic and there are often many scene changes in a video, by segmenting the video and separating different scenes, not only can multiple image semantics and multiple character relationships in the video be recognized, but the information obtained from the video is more comprehensive and accurate. Moreover, for each video segment, the interference of irrelevant information can be reduced, and the accuracy of the constructed index can be improved, so that the user's search intention can be accurately located in the subsequent search stage and the accuracy of the subsequent search can be improved.

[0254] In a fifth aspect, when performing processing such as character relationship analysis and semantic understanding, the method is based on representative frames extracted from each video segment. In this way, there is no need to process each frame image in the video segment separately, which can not only reduce noise information, but also reduce the amount of calculation, storage, etc., thus reducing resource overhead.

[0255] 2. Search stage

[0256] Exemplarily, Figure 10 is a flowchart of another visual media search method provided by an embodiment of the present application. As Figure 10 shown, the search stage of this method includes:

[0257] S201. In response to the operation of the user inputting a search text, the gallery APP sends the search text to the search module.

[0258] Referring to Figure 7 , for example, the user can enter the search text "traveled with family last year" in the search box. After receiving the text input by the user, the gallery APP sends the text to the search module.

[0259] S202. After receiving the search text, the search module sends a language understanding request to the natural language understanding module.

[0260] A language understanding request is used to request natural language understanding of the search text. Optionally, the search text can be carried in the language understanding request.

[0261] S203. In response to the language understanding request, the natural language understanding module performs natural language understanding on the search text, and identifies semantic entities, person relationships, category information, etc. included in the search text.

[0262] Optionally, the natural language understanding module can identify semantic entities, person relationships, and category information based on a natural language understanding model and / or named entity recognition (NER) technology.

[0263] In this embodiment, when performing natural language understanding on the search text "traveled with family last year", the semantic entity "last year" and the person relationship "family" can be identified.

[0264] S204. The natural language understanding module returns the identified semantic entities, person relationships, and category information to the search module.

[0265] S205. The search module sends a text semantic understanding request to the multimodal understanding module.

[0266] The text semantic understanding request is used to request semantic understanding of the search text.

[0267] Optionally, the search text can be carried in the text semantic understanding request.

[0268] S206. In response to the text semantic understanding request, the multimodal understanding module performs semantic understanding on the search text to obtain a text semantic vector.

[0269] Specifically, the multimodal understanding module can include a pre-trained text semantic understanding model. By inputting the search text into the text semantic understanding model, the corresponding semantic vector can be obtained. In this embodiment, the semantic vector corresponding to the text is referred to as the text semantic vector.

[0270] S207. The multimodal understanding module returns the text semantic vector to the search module.

[0271] S208. The search module performs multi-way recall in the inverted index library based on the semantic entities, person relationships, category information, and text semantic vector, and obtains multi-way recall results.

[0272] Specifically, the search module can perform semantic entity recall (also known as entity recall) in the inverted index library based on the semantic entity, that is, search for video segments or images matching the semantic entity in the inverted index library to obtain the semantic entity recall result; perform person relationship recall in the inverted index library based on the person relationship, that is, search for video segments or images whose person relationship information matches the person relationship in the search text to obtain the person relationship recall result; perform category information recall in the inverted index library based on the category information, that is, search for video segments or images whose labels match the category information in the inverted index library to obtain the category information recall result; perform semantic vector recall (abbreviated as vector recall) in the inverted index library based on the text semantic vector, that is, search for video segments or images whose visual semantic vectors match the text semantic vector in the inverted index library to obtain the vector recall result.

[0273] It should be noted that the above "matching" can be identity or the similarity meets the preset conditions.

[0274] Taking vector recall as an example, in some embodiments, the vector similarity between the text semantic vector and the visual semantic vector corresponding to the image or video segment in the inverted index library can be calculated. Optionally, the image or video segment with a vector similarity greater than the preset threshold can be selected for recall. Optionally, the obtained multiple vector similarities can also be sorted, and the top N indexes with higher vector similarities can be selected as the vector recall result. Here, N is an integer greater than 0. Exemplarily, N can be the number of preset vector recall results, such as 5, 8, or 10, etc.

[0275] The vector similarity refers to the degree of similarity between two vectors, which can be calculated by various methods. Exemplarily, the similarity degree can be determined by calculating the cosine similarity of the two vectors, or other methods can also be used. This application does not make any limitations in this regard.

[0276] Optionally, if the hit object is an image, the results of multi-channel recall can be represented by the ID of the image; if the hit object is a video segment, the results of multi-channel recall can be represented by the ID of the video segment and the ID of the video to which it belongs.

[0277] Optionally, when recalling, the search module can relax the recall conditions to a certain extent, that is, appropriately expand the search scope to make the recall results more comprehensive.

[0278] S209. The search module filters the multi-channel recall results according to the search text.

[0279] As described above, when performing vector recall, the recall conditions may be relaxed, so the recall results may contain many results that do not match the search statement. In addition, when performing multi-way recall, each way of recall may not consider the recall conditions of other ways of recall, so the recall results may also contain many results that do not match the search statement. In view of this, the search module can perform time filtering, spatial filtering, and person relationship filtering on the multi-way recall results to improve the accuracy of the final search results.

[0280] Specifically, the search module can filter out the results in the multi-way recall results that do not match the time included in the search text. For example, if the time in the search text is "last year", the results in the multi-way recall results with an acquisition time that does not overlap with "last year" can be filtered out. For example, the recall result with an acquisition time of January 1st this year can be filtered out.

[0281] The search module can filter out the results in the multi-way recall results that do not match the location included in the search text. For example, if the location in the search text is "Haidian District, Beijing", the recall result with an acquisition location of "Xicheng District, Beijing" in the multi-way recall results can be filtered out.

[0282] The search module can filter out the results in the multi-way recall results that do not match the person relationship included in the search text. For example, if the person relationship in the search text is "me and my son", the recall results with person relationship information including only "me" and the recall results with person relationship information including only "son" in the multi-way recall results can be filtered out.

[0283] S210. The search module sorts the multi-way recall results to obtain search results.

[0284] In some embodiments, the intersection or union of the multi-way recall results can be sorted.

[0285] Exemplarily, each way of recall results may include a scoring result. This scoring result represents the matching degree between the recall result and the search text, and the multi-way recall results can be sorted by weighted fusion according to the scoring result. Optionally, the multi-way recall results can be sorted in descending order of the scoring result. The higher the scoring result, the more it indicates a match with the search text.

[0286] S211. The search module returns the search results to the gallery APP.

[0287] The search module returns the sorted search results to the gallery APP.

[0288] S212. The gallery APP displays the search results.

[0289] The search results can be displayed, for example, as described above Figure 5The search result interface 501 shown in FIG. (b) therein, or Figure 6 the search result interface 601 shown in FIG. (b) therein.

[0290] In this embodiment, natural language understanding is performed on the search content input by the user to identify semantic entities, person relationships, category information, etc. therein, and semantic understanding of the search content is performed to obtain a text semantic vector. Then, multi-channel recall is performed based on the semantic entity, person relationship, category information, and text semantic vector. In this way, matching is performed with the search text from multiple dimensions, that is, all-round search from multiple dimensions, greatly improving the matching degree between the search result and the search text and improving the user experience. Moreover, sorting the multi-channel recall results enables the user to preferentially see visual media with a high matching degree with the search text, improving the user experience.

[0291] The video segmentation process will be further described below with reference to the accompanying drawings.

[0292] Exemplarily, Figure 11 is a flowchart showing a video segmentation process provided by an embodiment of the present application. As Figure 11 shown, in the above step S104, "performing segmentation processing on the incremental video to obtain multiple video segments and their attribute information, and extracting representative frames from each video segment" includes the following steps. The execution subject of the following steps is the video segmentation module, which will not be elaborated.

[0293] S301. Perform frame splitting processing on the incremental video to obtain multiple frame images.

[0294] S302. Based on the labels and parameters of each frame image, determine the segmentation score of the frame image, where the segmentation score characterizes the degree of change between the frame image and its adjacent frame image.

[0295] It can be understood that when playing a video, the content displayed by the video changes continuously as the frame images are played in sequence, but the degree of change is different. Therefore, similar (i.e., with a small degree of change) frame images can be classified into one video segment.

[0296] In some embodiments, the segmentation score calculated through the frame label of the frame image and the frame parameter of the frame image can characterize the degree of change between a frame image and its adjacent frame image, and then the video is segmented based on the segmentation score. The frame parameter is used to characterize the display characteristics of the frame image and can also be called an image parameter. Exemplarily, the frame parameter of the frame image can include the jitter degree, clarity, pixel value, etc. of the frame image, and the present application does not limit this.

[0297] Among them, the jitter of a frame image refers to the phenomenon that the content shown in the frame image jitters or shakes during video playback. Exemplarily, when a user holds a mobile phone to take a picture and moves the phone when hoping to capture another scene, obvious jitter may occur. The jitter degree is used to characterize the jitter degree of the frame image.

[0298] The clarity of a frame image refers to the clarity of each detail texture and its boundary in the frame image.

[0299] The pixel value of a frame image can represent the brightness of the frame image.

[0300] In some embodiments, based on the frame label and frame parameters of the frame image, the segmentation score of the frame image can be calculated by combining the following formula (1), and then it can be determined whether to determine it as the start frame or the end frame of a video segment according to the segmentation score of the frame image. Formula (1) is as follows:

[0301] y = α × frame A + β × frame B + γ × frame Y + δ × frame T (1)

[0302] Among them, y represents the segmentation score of the frame image, frame A represents the jitter degree score of the frame image, frame B represents the clarity change score of the frame image, frame Y represents the label change score of the frame image, frame T represents the pixel change score of the frame image, and α, β, γ, and δ respectively represent frame A , frame B , frame Y and frame T 's coefficients. In some embodiments, α, β, γ, and δ can be values preset manually according to the influence degrees of the jitter degree, clarity change, label change, and pixel change of the frame image on the segmentation score of the frame image.

[0303] It should be noted that the process of determining the segmentation score of the frame image based on the frame label and multiple frame parameters above is only an example. It can also be based on one or more of the frame label and multiple frame parameters to determine the segmentation score of the frame image, and this application does not make any limitations in this regard.

[0304] In some embodiments, video jitter detection methods such as the optical flow method, feature point matching method based on image displacement, and based on image gray distribution characteristics can be used to detect the jitter degree of multiple frame images in a video. Then, based on the correspondence relationship between the preset jitter degree range and the jitter degree score, the jitter degree score of the frame image can be determined.

[0305] In some embodiments, a clarity detection tool may be used to determine the clarity of multiple frame images in a video, and then the clarity of the frame images is compared with the clarity of adjacent frame images to determine the clarity change value of the frame images. Subsequently, based on the correspondence between the preset clarity change value range and the clarity change score, the clarity change score of the frame images is determined. Exemplarily, the video quality detection tool may be open-source software such as FFmpeg and Video Quality Measurement Tool.

[0306] In some embodiments, the frame label of a frame image may be compared with the frame label of an adjacent frame image, and the comprehensive label change of the frame image is determined based on the quantity change and content change of the frame labels of the frame image. Exemplarily, the union and intersection of the frame label of a frame image and the frame label of its adjacent frame image may be calculated. The quantity change of the frame labels in the union can reflect the quantity change of the frame labels of the frame image, and the quantity change of the frame labels in the intersection can reflect the content change of the frame labels. When there is a change in the quantity of the frame labels in the union or the quantity of the frame labels in the intersection, the label change score of the frame image can be determined according to the correspondence between the preset quantity change range of the frame labels and the label change score.

[0307] In some embodiments, a pixel detection tool may be used to determine the pixels of multiple frame images in a video, and then the pixels of the frame images are compared with the pixels of adjacent frame images to determine the pixel change value of the frame images. Subsequently, based on the correspondence between the preset pixel change value range and the pixel change score, the pixel change score of the frame images is determined. Exemplarily, the pixel detection tool may be plugins such as PixelStick, MeasureIt, and Guides.

[0308] S303. Segment the incremental video according to the segment scores of each frame image.

[0309] In some embodiments, the degree of change in the content shown by two adjacent frame images in the video may be positively correlated with the magnitude of the segment score of the frame image, that is, the larger the segment score of the frame image, the more obvious the degree of change in the content shown compared with the adjacent frame image.

[0310] Exemplarily, a frame image may be compared with the previous frame image, and a segment score threshold is preset, and the frame image whose segment score exceeds the segment score threshold is determined as the starting frame of a video segment. As shown in Formula 1 above, the higher the jitter degree of the frame image, the higher the corresponding value; the greater the clarity change of the frame image compared with the previous frame image, the higher the frame A corresponding value; the greater the clarity change of the frame image compared with the previous frame image, the higher the frame BThe higher the corresponding value; compared with the previous frame image, the greater the change in the label of this frame image, frame Y The higher the corresponding value; compared with the previous frame image, the greater the change in the label of this frame image, frame Y The higher the corresponding value; compared with the previous frame image, the greater the change in the pixels of this frame image, frame T The higher the corresponding value.

[0311] Exemplarily, a frame image can be compared with its next frame image, and a segmentation score threshold can be preset, and the frame image whose segmentation score exceeds the segmentation score threshold is determined as the end frame of a video segment.

[0312] In some other embodiments, the degree of change in the content shown by two adjacent frame images in a video can also be negatively correlated with the size of the segmentation score of the frame image, that is, the smaller the segmentation score of the frame image, the more obvious the degree of change in the content it shows compared with the adjacent frame image, that is, the lower the segmentation score of the frame image, the greater the possibility of determining it as the start frame or the end frame of a video segment.

[0313] Exemplarily, a frame image can be compared with its previous frame image, and a segmentation score threshold can be preset, and the frame image whose segmentation score is lower than the segmentation score threshold is determined as the start frame of a video segment. As shown in the above formula (1), the higher the jitter degree of this frame image, frame A The lower the corresponding value; compared with the previous frame image, the greater the change in the clarity of this frame image, frame B The lower the corresponding value; compared with the previous frame image, the greater the change in the label of this frame image, frame Y The lower the corresponding value; compared with the previous frame image, the greater the change in the label of this frame image, frame Y The lower the corresponding value; compared with the previous frame image, the greater the change in the pixels of this frame image, frame T The lower the corresponding value.

[0314] S304. Determine the corresponding representative frames from each video segment.

[0315] In some embodiments, the representative frame can be the start frame, end frame, middle frame or random frame in a video segment.

[0316] Exemplarily, assuming that a video segment contains 9 frame images, the first frame image (start frame), the ninth frame image (end frame) or the fifth frame image (middle frame) among these 9 frame images can be used as the representative frame of this video segment according to the time sequence, or a frame image (random frame) can be randomly selected from these 9 frame images as the representative frame of this video segment.

[0317] In some embodiments, the optimal frame can also be determined from a video segment as the representative frame based on preset rules related to the frame parameters and frame tags of the frame images.

[0318] Exemplarily, the score of a frame image can be calculated based on the frame parameters and frame tags of the frame image. The more tags the frame image has, the higher its score; the lower the jitter degree of the frame image, the higher its score; the higher the clarity of the frame image, the higher its score; and the lower the pixel change of the frame image compared to the pixels of the previous frame image, the higher its score. Finally, the frame image with the highest score (i.e., the optimal frame) in a video segment can be determined as the representative frame.

[0319] In this implementation manner, considering various aspects such as true parameters and frame tags in the video segment, the optimal frame is determined from the video segment as the representative frame, which reduces the probability of selecting an interfering frame as the representative frame, reduces the perturbation of irrelevant information, improves the accuracy of indexing, and thus improves the accuracy of video search.

[0320] Next, in combination with Figure 12 the process of determining the representative frame in the video will be introduced in detail.

[0321] As Figure 12 shown, the video includes 120 frame images. One frame image is compared with the previous frame image, and the segmented scores corresponding to these 120 frame images are calculated in combination with the above formula (1). The frame images whose segmented scores exceed the segmented score threshold are used as the starting frames of a video segment.

[0322] In some embodiments, these 120 frame images are divided into three video segments, namely video segment 1, video segment 2, and video segment 3, and the middle frame of each video segment is determined as its corresponding representative frame. Video segment 1 includes 40 frame images, video segment 2 includes 51 frame images, and video segment 3 includes 29 frame images. That is, the 20th frame image in video segment 1 is used as the representative frame corresponding to video segment 1, the 26th frame image in video segment 2 is used as the representative frame corresponding to video segment 2, and the 15th frame image in video segment 3 is used as the representative frame corresponding to video segment 3.

[0323] It should be noted that the number of frame images in video segment 1 is 40, which is an even number. Both the 20th frame image and the 21st frame image are middle frames, and any one of the middle frames can be determined as the representative frame corresponding to video segment 1. This application does not make any limitations in this regard.

[0324] Next, in combination with the accompanying drawings, the specific process of analyzing the relationship between characters will be described.

[0325] Exemplarily, Figure 13 is a schematic flowchart of the process of analyzing the relationship between characters provided by an embodiment of this application. Figure 14A logic schematic diagram of a person relationship analysis provided by an embodiment of this application. Please also refer to Figure 13 and Figure 14 In this embodiment, the specific process of identifying the social circle and the type of social circle in the above step "S113. The person relationship analysis module responds to a person relationship analysis request, and based on the face clustering result and the face image set in the CV database, performs person relationship analysis to obtain the person relationship information of the incremental face image and the face representative frame" is mainly described. The execution subject of the following steps is the person relationship analysis module, which will not be elaborated.

[0326] The images involved in the following embodiments are all face images in the face image set. This image may be an incremental image or a representative frame. For the sake of convenience of description, hereinafter they are all referred to as images or images containing people, and no further detailed distinction will be made. It can be understood that some attribute information of the images involved in the following description, such as the acquisition location of the image, the acquisition time of the image, etc., for the representative frame, is the attribute information corresponding to the video to which it belongs. No further detailed distinction will be made hereinafter either.

[0327] S401. Generate a person-image heterogeneous graph according to the face clustering result. The person-image heterogeneous graph is used to represent the corresponding relationship between a person and an image containing the person.

[0328] Specifically, each person in the face clustering result can be used as a person node, each image including a person can be used as an image node, and the image node including a certain face is connected to the person node corresponding to the face by an edge to obtain a person-image heterogeneous graph. For example, if face 1 is included in image A, according to the face clustering result, the person corresponding to this face 1 is a, then the node of image A is connected to the node of person a by an edge; the same applies to other images and other people in the face clustering result. In this way, a person-image heterogeneous graph can be obtained. That is to say, the person-image heterogeneous graph contains person nodes and image nodes. The edge between a certain person node and an image node indicates that the image contains the person, or in other words, the edge between a certain person node and an image node indicates that the person belongs to the image.

[0329] Optionally, the person-image heterogeneous graph may also include the node information of the person node and the image node. Specifically, the information of the person node is also the information of the person corresponding to the person node, including but not limited to person ID, person gender, and person age. The information of the image node is also the attribute information of the image corresponding to the image node, including but not limited to image ID, acquisition time of the image, acquisition location of the image, and the camera used for image shooting (such as whether it is taken by a front camera).

[0330] Exemplarily, Figure 15 A person-image heterogeneous graph provided by an embodiment of this application.Figure 15 In it, the illustration with a circular border represents a person node, and the illustration with a square border represents an image node. It can be understood that Figure 15 only as an example, used to illustrate the composition structure of the person-image heterogeneous graph, etc., not as a limitation, nor does it represent actual data.

[0331] It should be noted that Figure 15 is a concrete representation of the person-image heterogeneous graph. In practical applications, there can be various forms of representation of the person-image heterogeneous graph. The embodiments of this application do not make any limitations on the specific form of representation of the person-image heterogeneous graph, as long as its content can be reflected.

[0332] In a specific embodiment, the person-image heterogeneous graph can also be represented by a list. Specifically, the table corresponding to the person-image heterogeneous graph can include a person node list, an image list, a connection matrix list, etc. The person node list is used to save the information of each person node. The image list is used to save the information of each image node. The connection matrix list is used to save the connection relationship between the person node and the image node, where the connection relationship can be represented in the form of a matrix. The connection matrix list reflects the information of the edges in the person-image heterogeneous graph, representing the people included in each image, or the images to which each person belongs.

[0333] Exemplarily, the person node list can be as shown in Table 3, the image node list can be as shown in Table 4, and the connection matrix list can be as shown in Table 5.

[0334] Table 3

[0335]

[0336] Table 4

[0337]

[0338] Table 5

[0339] Serial number Image node Matrix of person nodes connected to the image node 1 Image node A [Person node a, Person node b, Person node c] 2 Image node B [Person node b, Person node c] 3 Image node C [Person node a] … … … n Image node N [Person node a, Person node n]

[0340] It should be noted that the above Tables 3, 4, and 5 are only for examples, used to exemplify the presentation form of the person-image heterogeneous graph and the content of node information, etc., do not represent actual data, nor do they constitute any limitation to this application. In actual use, Tables 3, 4, and 5 can include more or less content.

[0341] S402. Determine the co-occurrence image set according to the person-image heterogeneous graph.

[0342] Among them, a co-occurrence image refers to an image in which there is a co-occurrence relationship of characters. The existence of a co-occurrence relationship of characters means that two or more characters appear in the same image. That is to say, a co-occurrence image refers to an image containing at least two characters. For example, if an image A includes character a, character b, and character c, then image A is determined to be a co-occurrence image. If an image B only includes character a, then image B is determined not to be a co-occurrence image.

[0343] Specifically, according to the person-image heterogeneous graph, the number of person nodes connected to an image node can be determined. If a certain image node is connected to two or more person nodes, the image corresponding to the image node is a co-occurrence image. Generate a co-occurrence image set from the determined multiple co-occurrence images. It can be understood that in other embodiments, the determined multiple co-occurrence images can also be presented in other forms (such as a list) instead of in a combined form. The embodiments of the present application do not make any limitations in this regard.

[0344] In addition, for every two characters in a co-occurrence image, the co-occurrence image can be called the co-occurrence image of these two characters. For example, if image A includes character a, character b, and character c, then image A is the co-occurrence image of character a and character b, and also the co-occurrence image of character a and character c, and also the co-occurrence image of character b and character c.

[0345] For the convenience of description, in subsequent embodiments, the co-occurrence image of character a and character b is simply referred to as the ab co-occurrence image of characters, the co-occurrence image of character a and character c is called the ac co-occurrence image of characters, and the co-occurrence image of character b and character c is called the bc co-occurrence image of characters, and so on. The co-occurrence image of any two characters is also called the target co-occurrence image.

[0346] According to the person-image heterogeneous graph, the co-occurrence image of every two characters can also be determined. Specifically, for any two characters (taking character a and character b as an example), according to the connection matrix list, the matrix in the person node matrix that includes both the person node corresponding to character a and the person node corresponding to character b can be determined, and the image nodes corresponding to these matrices can be determined, so as to determine the co-occurrence image of character a and character b. Optionally, the co-occurrence image of every two characters can be represented in the form of a matrix. Taking the connection matrix list shown in Table 5 as an example, according to the person node matrix, it is determined that the matrix numbered 1 includes character a and character b, and the image node corresponding to this matrix is image node A. Therefore, it is determined that the image A corresponding to image node A is the ab co-occurrence image of characters. Referring to this method, all ab co-occurrence images of characters can be obtained, and the co-occurrence images of any other two characters can also be obtained.

[0347] S403. Determine the intimacy between every two characters in the face clustering result according to the attribute information of each image in the co-occurrence image set.

[0348] Specifically, the intimacy between two characters is used to characterize the intimacy of the co-occurrence relationship between these two characters, that is, to characterize the degree of association between the two characters.

[0349] It can be understood that the intimacy between two characters is a value greater than or equal to 0. The greater the intimacy, the closer the two characters are. If the intimacy is 0, it means that there is no social relationship between the two characters.

[0350] Specifically, for the characters included in the images in the co-occurrence image set, the intimacy between these characters in pairs can be determined according to the attribute information of each image in the co-occurrence image set. For the characters included in the images that do not belong to the co-occurrence image set (i.e., non-co-occurrence images), there is no co-occurrence relationship between these characters and other characters. Therefore, the intimacy between these characters and other characters is determined to be 0. In this way, the intimacy between every two characters among all the characters in the face clustering result can be obtained.

[0351] As a possible implementation, the PMI between every two characters among all the characters included in the images in the co-occurrence image set can be determined according to information such as the acquisition time and acquisition location of each image in the co-occurrence image set. Then, the intimacy between the two characters is determined according to the PMI between the two characters. Among them, the attribute information of the co-occurrence image can be obtained through the node information corresponding to the image node in the character-image heterogeneous graph. Specifically, it can be obtained from the image node list shown in Table 4. Of course, the corresponding attribute information can also be directly searched from the CV database according to the ID of the co-occurrence image. The embodiments of the present application do not make any limitations on this.

[0352] The specific implementation process of step S403 will be further described in detail in subsequent embodiments.

[0353] S404. Simplify the character-image heterogeneous graph into a character connection graph according to the intimacy between every two characters.

[0354] The character connection graph is used to characterize the social relationship and the corresponding intimacy degree between characters. The character connection graph is a homogeneous graph.

[0355] Specifically, based on the character-graph heterogeneous graph, the image nodes are removed. Then, two character nodes with non-zero intimacy are connected by an edge, and the intimacy is used as the weight of this edge. In this way, the character connection graph can be obtained.

[0356] In one embodiment, for Figure 15 the character-image heterogeneous graph shown, the obtained character connection graph after simplification can be as Figure 16 shown.

[0357] Similarly, it can be understood that Figure 16It is a concrete representation of a person connection graph. In practical applications, there can be various forms of expression for the person connection graph. For example, the person connection graph can be represented by a list. Specifically, based on the above Table 3, Table 4, and Table 5, Table 4 can be deleted, and information such as the gender and age of the people in Table 3 can be deleted to obtain a simplified list of person nodes. Then, the image nodes in Table 5 and the person nodes that have no co-occurrence relationship with other people can be deleted to obtain a simplified list of connection matrices. The simplified list of person nodes and the list of connection matrices are used as the person connection graph. The embodiments of the present application do not make any limitation on the specific form of expression of the person connection graph, as long as its content can be reflected.

[0358] S405. Identify the central person according to the person-image heterogeneous graph.

[0359] The central person refers to the person in the face clustering result who is in the social center. Generally, the central person has more and closer social relationships with other people. It can be understood that the number of central persons can be one or more, and can be specifically set according to actual needs. The central person can be the owner of the electronic device or a non-owner, such as the family member of the owner, etc.

[0360] Research has found that the images containing the central person may be generated at multiple times and multiple locations. Therefore, the distribution of the images containing the central person in terms of time and location will be relatively uniform. That is to say, the divergence of the time distribution and location distribution of the images containing the central person is small. Compared with non-central persons, the images containing the central person generally will not be generated only at one or several specific times, and / or will not be generated only at one or several specific locations. For example, when the central person is the owner himself / herself, the owner may take pictures with the mobile phone at any time to generate images containing himself / herself. For a non-owner, such as a friend of the owner, the friend may only take pictures with the owner's mobile phone during a period of time when walking with the owner to generate images containing the friend. Therefore, the distribution of the images containing the owner himself / herself is more uniform than the distribution of the images containing the owner's friend.

[0361] Based on this, in the present application, the central person can be determined according to the divergence of the time distribution and location distribution of the images of each person in the person-image heterogeneous graph. Specifically, among the people included in the face clustering result, the person whose divergence of the time distribution and location distribution of the corresponding images meets the preset conditions can be determined as the central person. Among them, the image of a certain person is also called the image corresponding to that person, which refers to the image containing that person. For example, the image of person a is also called the image corresponding to person a, which refers to the image containing person a, the image of person b is also called the image corresponding to person b, which refers to the image containing person b, and so on.

[0362] Optionally, the temporal divergence of the images of each person can be characterized by the temporal KL divergence, and the spatial divergence of the images of each person can be characterized by the spatial KL divergence.

[0363] The specific implementation process of step S405 will be further described in detail in the following embodiments.

[0364] S406. Remove the central person node in the person connection graph to obtain a de-centered person connection graph.

[0365] Specifically, based on the person connection graph, remove the node corresponding to the identified central person, and remove the edges connecting the central person node and other person nodes to obtain a de-centered person connection graph.

[0366] In one embodiment, taking the Figure 16 shown person connection graph as an example, if the central person is the person corresponding to the person node filled with solid black in Figure 16 , then based on this person connection graph, after removing the central person node, the obtained de-centered person connection graph can be as shown in Figure 17 .

[0367] Similar to the person-image heterogeneous graph and the person connection graph, Figure 17 only as a concrete representation of the de-centered person connection Figure 1 type, in practical applications, the representation form of the de-centered person connection graph can be various, and the embodiments of the present application do not make any limitations and will not be elaborated herein.

[0368] S407. Perform community detection based on the de-centered person connection graph to obtain at least one social circle.

[0369] Among them, the result of community detection is multiple communities, and each community corresponds to a social circle. Each social circle includes multiple (at least two) persons.

[0370] In one embodiment, community detection can be performed based on the Louvain algorithm to obtain at least one social circle. Specifically, it can include the following steps:

[0371] 1) Use the de-centered person connection graph as the target connection graph.

[0372] 2) Take each community in the target connection graph as a node, and perform a modularity optimization operation on each node in the target connection graph to obtain an updated connection graph.

[0373] Specifically, the modularity optimization operation includes: for any node x, attempt to add node x to each adjacent community respectively, and calculate the corresponding modularity gain; if the maximum modularity gain among all modularity gains is negative, then do not change node x, that is, do not add node x to any adjacent community; if the maximum modularity gain among the modularity gains of adjacent communities is not negative, add node x to the adjacent community corresponding to the maximum modularity gain.

[0374] Among them, the adjacent community of node x refers to the community where the nodes adjacent to node x are located. It should be noted that for a single node (i.e., a node that has not joined a community), it itself is a community. In other words, the community where a single node is located only includes one node, which is the node itself.

[0375] After traversing all nodes in the target connection graph through the modularity optimization operation, an updated connection graph is obtained.

[0376] Specifically, when adding node x to any adjacent community (taking adjacent community C as an example), the calculation process of the corresponding modularity gain can be as follows:

[0377] a. Calculate the modularity of the target connection graph when node x is added to adjacent community C to obtain the first modularity.

[0378] b. Calculate the modularity of the target connection graph when node x is not added to adjacent community C to obtain the second modularity.

[0379] c. Calculate the difference between the first modularity and the second modularity to obtain the modularity gain corresponding to adjacent community C.

[0380] Among them, in steps a and b, the first modularity and the second modularity can be calculated with reference to formula (2):

[0381]

[0382] In formula (2), Q represents the modularity of the target connection graph. W represents the sum of the weights of all edges in the target connection graph. Both i and j are the node numbers in the target connection graph, i represents node i, and j represents node j. s i represents the sum of the weights of all edges connected to node i in the target connection graph. s j represents the sum of the weights of all edges connected to node j in the target connection graph. A i,j represents the weight of the edge between node i and node j. k i represents the sum of the weights of all edges pointing to node i. δ(C i ,C j ) is an indicator function, C i represents the community where node i is located, C jdenotes the community where node j is located. When node i and node j belong to the same community (i.e., C i is the same as C j ), the value of δ(C i , C j ) is 1; when node i and node j do not belong to the same community (i.e., C i is not the same as C j ), the value of δ(C i , C j ) is 0.

[0383] 3) Take the updated connection graph as the target connection graph, take each community in the updated connection graph as a node respectively, and return to execute step 2) until the nodes included in all communities no longer change, obtaining at least one community.

[0384] That is to say, when executing step 2) for the first time, the target connection graph is the decentralized person connection graph, and each node in the decentralized person connection graph is a single node, that is, a node that has not joined the community, so each node is taken as a community. When executing step 2) for the second time and subsequent times, the target connection graph is the updated connection graph, and the updated connection graph includes both single nodes and discovered communities. Therefore, taking the communities in the updated connection graph as nodes, perform modularity optimization operations respectively. During the process of performing the modularity optimization operation, for a single node in the updated connection graph, the community it belongs to only includes itself. In this way, iterative implementation is achieved, changing the communities to which nodes belong, optimizing the modularity of each community until the modularity gain of each community no longer increases and the nodes included in each community no longer change.

[0385] In this embodiment, community discovery is performed based on the Louvain algorithm, with low time complexity during calculation and stable community division results, making the generated social circle results stable and accurate. Of course, in some embodiments, other algorithms can also be used for community discovery to obtain a social circle. For example, community discovery can be performed based on the infomap algorithm. The specific method for community discovery in the embodiments of this application is not limited in any way.

[0386] 4) Determine at least one social circle according to at least one community.

[0387] In one embodiment, after performing community discovery based on the Louvain algorithm and obtaining at least one community, the set of people corresponding to all nodes in each community can be directly determined as a social circle, and in this way, the social circles corresponding to each community can be obtained.

[0388] In another embodiment, after performing community discovery based on the Louvain algorithm, the obtained communities can also be filtered to obtain the final social circle. Specifically, it can include:

[0389] For any community, calculate the sum of the weights of the edges connecting each person node in the community to other person nodes in the community; remove the N person nodes with the smallest sum of weights from the community to obtain a de-marginalized community; determine the set of the persons corresponding to all the nodes in the de-marginalized community as a social circle.

[0390] Among them, the N person nodes with the smallest sum of weights can be called marginal person nodes. That is to say, remove the marginal person nodes in the community to obtain a de-marginalized community.

[0391] For example, according to the above steps 1) to 3), it is determined that a certain community C obtained includes person node A1, person node A2, person node A3... person node An. Calculate the sum of the weights of the edges connecting person node A1 to other person nodes (person node A2, person node A3... person node An) in community C; then calculate the sum of the weights of the edges connecting person node A2 to other person nodes (person node A1, person node A3... person node An) in community C... and so on, calculate the sum of the weights of the edges connecting each person node to other person nodes. Remove the N person nodes with the smallest sum of weights from community C. For example, if N = 1, then remove the 1 person node with the smallest sum of weights from community C to obtain marginal community C. Determine the set of the persons corresponding to all the nodes in marginal community C as a social circle.

[0392] Exemplarily, Figure 18 FIG. is a schematic diagram of a scenario for removing marginal person nodes provided by an embodiment of the present application. Figure 18 In, the black dots represent person nodes, and the connections between the person nodes are edges. The thickness of the edge represents the magnitude of the weight of the edge. The thicker the edge, the greater the weight of the edge, and vice versa. As Figure 18 shown, through community discovery, community C is obtained. Among them, community C includes person node E. After calculation, the weight of the edge connecting person node E to other nodes in community C is the smallest, and N = 1. Therefore, remove the person node E from community C to obtain de-marginalized community C. Determine the set of the persons corresponding to all the nodes in de-marginalized community C as a social circle.

[0393] Process each community according to the above process to obtain the social circle corresponding to each community.

[0394] As described above, the weight of an edge is the intimacy between two characters. Therefore, the sum of the weights of the edges connecting a character node to other character nodes represents the number of connections between this character node and other character nodes in the community, as well as the magnitudes of the weights of each edge. A smaller sum of weights indicates that this character node has fewer connections with other characters in the community, and / or the weights of the edges connecting this node to other nodes are smaller. This further indicates that the character corresponding to this character node has fewer associations with other characters in the social circle, and / or lower intimacy with other characters. That is, this character is a character with a lower connectivity, and is considered an edge character for this social circle. In this embodiment, edge characters are filtered out, and only characters with a higher character connectivity are included in the social circle, making the identified social circle more in line with the user's actual social situation and further improving the user experience.

[0395] Optionally, the result of community discovery to obtain social circles can be represented in the form of a connection graph, that is, a social circle graph is generated. The social circle graph is used to represent the characters included in each social circle, as well as the social relationships and corresponding intimacies between the characters.

[0396] It can be understood that after community discovery to obtain social circles, the central character node can also be added to the social circle graph according to the character connection graph to obtain a character - social circle graph.

[0397] In one embodiment, the character - social circle graph can be as Figure 19 shown. Figure 19 In [the figure], the black dots represent character nodes, and the lines connecting the character nodes are edges. The thickness of the edge represents the magnitude of the weight of the edge. The thicker the edge, the greater the weight of the edge, and vice versa. It can be seen that the central character node is connected to multiple character nodes, and some of the edges have larger weights. Among them, Figure 19 the social circles shown in [the figure] include three: Social Circle 1, Social Circle 2, and Social Circle 3.

[0398] It should be noted that the above process of community discovery can ignore node information and be based only on the weights of nodes and edges. Therefore, in the character connection graph generated in step S404, the information of the nodes may not be included, and only the nodes, edges, and the weights of the edges (i.e., intimacies) are included. This can further simplify the structure of the character connection graph and improve the algorithm operation efficiency.

[0399] In addition, in this embodiment, community discovery is performed based on the decentralized person connection graph to obtain social circles. The reason is as follows: Generally speaking, in a person connection graph, the central person node has many edges connected to other nodes, and the intimacy between the central person and other persons is relatively high. Therefore, the weight of the edges connected to the central person node will be very high. If the central person node is included in community discovery, it is very easy to divide the central person node into a certain community in a certain round of iteration. In this way, in the next round of iteration, when the community where the central person node is located is used as a node, there are many edges with other nodes, and / or the sum of the weights of the edges pointing to the community where the central person node is located will be very high. Therefore, other points that do not belong to this community will be divided into this community, resulting in inaccurate social circles finally generated. Therefore, in this embodiment, community discovery is performed based on the decentralized person connection graph to obtain social circles, so as to improve the accuracy of community discovery, and further improve the accuracy of the obtained social circles and improve the user experience.

[0400] S408. Determine the types of each social circle according to the central person and the person-image heterogeneous graph. The type of a social circle represents the type of relationship between the persons in the social circle and the central person.

[0401] As a possible implementation manner, the types of each social circle can be determined based on a graph neural network model and a multilayer perceptrons (MLP) model. Specifically, the following process can be included:

[0402] 1) Input the person-image heterogeneous graph into the pre-trained graph neural network model to obtain the feature vectors of each person node and image node.

[0403] The graph neural network model is used to extract the feature vectors of the nodes. Optionally, an initial graph neural network model can be established first. Then, the initial graph neural network model is trained with the person-image heterogeneous graph samples of the labeled data to obtain the graph neural network model. When determining the types of each social circle, the person-image heterogeneous graph is input into this graph neural network model to obtain the feature vectors of each person node and image node.

[0404] Among them, the person-image heterogeneous graph includes the information of each person node and the information of the image node. As described above, the information of the person node can include the person's gender, age, etc. The feature vectors of the person nodes extracted by the graph neural network model can include the person's gender feature vector, age feature vector, etc. Optionally, the value of the person's gender feature vector is 0 or 1. Optionally, a value of 0 indicates male gender, and a value of 1 indicates female gender. The value range of the person's age feature vector can be [0, 100], and the value of the person's age can be an integer.

[0405] As described above, the information of the image node may include the acquisition time of the image, the acquisition location of the image, whether it is taken by the front camera, etc. The feature vectors of the image nodes extracted by the graph neural network model may include: acquisition time feature vector, acquisition location feature vector, and whether front camera feature vector, etc. Among them, the acquisition time feature vector may include acquisition moment feature vector, hour information feature vector, month difference feature vector, and week feature vector, etc. Specifically, the hour information feature vector refers to the feature vector of the hour information in the acquisition moment. The value range of the hour information feature vector may be [0, 23], and the value of the hour information feature vector is an integer. The month difference feature vector refers to the feature vector of the difference between the month information in the acquisition moment and the month information in the current moment. The value range of the month difference feature vector may be [0, 100], and the value of the month difference feature vector is an integer. The week feature vector refers to the feature vector of the week information in the acquisition moment. The value range of the week feature vector may be [1, 7], and the value of the week feature vector is an integer. The acquisition location feature vector may include one-hot vectors defined such as home, company, scenic spot, etc. The value of the whether front camera feature vector may be 0 or 1. Optionally, a value of 0 indicates that it is not taken by the front camera, and a value of 1 indicates that it is taken by the front camera.

[0406] Optionally, the graph neural network model may be a relational data modeling with graph convolutional networks (R-GCN) model. The main idea of the R-GCN model for extracting feature vectors is that the feature vector of node i at the l+1 layer can be obtained by aggregating the feature vectors of neighbor nodes of different relationship types. To distinguish and learn the data of each relationship, neighbor nodes of different relationships are transformed with different transformation matrices (also called weight matrices).

[0407] In a specific embodiment, the calculation formula of the R-GCN model may be as follows in formula (3):

[0408]

[0409] Where, represents the feature vector of node i at the l+1 layer. represents the set of neighbor nodes whose relationship with node i is r. For the person-image heterogeneous graph, there is only one such relationship r, namely the person and image inclusion relationship. c i,r represents the regularization constant. Optionally, c i,r can be set according to requirements. For example, c i,r can take a value of represents the pair Perform determinant calculation. Represents the transformation matrix of the l-th layer of the relationship r. For neighbor nodes of the same relationship (i.e., edges of the same type), the same transformation matrix is used. Perform transformation. Optionally, It can be a linear transformation function.

[0410] 2) Input the feature vectors of each person node and image node into a pre-trained multi-layer perceptron model to obtain the relationship type between each person node and the central person node (abbreviated as the relationship type of the person node).

[0411] The multi-layer perceptron model is used to predict the type of relationship between each person node and the central person node. The types of relationships between person nodes and the central person node can include but are not limited to family members, colleagues, and classmates.

[0412] Optionally, an initial multi-layer perceptron model can be established first. Then, the initial multi-layer perceptron model is trained with the node feature vector samples of the labeled data to obtain the multi-layer perceptron model. When determining the type of a social circle, input the feature vectors of the person nodes and image nodes obtained in step 1) into this multi-layer perceptron model to obtain the relationship type between each person node and the central person node, that is, obtain the relationship type of each person node.

[0413] 3) Count the number of people corresponding to each type of relationship in each social circle respectively, and determine the type of the social circle as the relationship type with the largest number of people.

[0414] For example, for any social circle C, the relationship types of the people in this social circle C include "colleague" and "classmate". Then count the number of people with the relationship type of "colleague" in social circle C (for example, it is x1), and count the number of people with the relationship type of "classmate" in the social circle (for example, it is x2). If x1 is greater than x2, then determine "colleague" as the type of this social circle C.

[0415] 4) Modify the relationship types of the people (i.e., the nodes corresponding to the people to be corrected) in each social circle that are different from the type of the social circle to the type of the social circle.

[0416] For example, in the above example, if the determined type of social circle C is "colleague", then modify the relationship types of each person in social circle C whose relationship type is not "colleague" to "colleague". In this way, it is possible to correct the relationship types of the people with biased judgments in this social circle C, so that the relationships between the people and the central person in the determined social circle C are consistent and accurate, that is, improve the accuracy of person relationship recognition.

[0417] Exemplarily, Figure 20 ForFigure 17 Decentralized person connection diagram, schematic diagram of the recognition result obtained by identifying the social circle and the type of social circle. As Figure 20 shown, in this embodiment, two social circles are recognized. The types of these two social circles are family members and colleagues respectively. Based on this, the relationship type between each person in the "family member" social circle and the central person (shown as a solid black fill in the figure) is corrected to "family member" (shown as a hatched pattern fill in the figure), and the relationship type between each person in the "colleague" social circle and the central person is corrected to "colleague" (shown as a black dot pattern fill in the figure).

[0418] In this embodiment, according to the face clustering result, a person-image heterogeneous graph is generated, and based on the person-image heterogeneous graph, the intimacy between two persons is determined. Then, based on the intimacy and the person-image heterogeneous graph, a person connection diagram is generated, and the central person is identified based on the person-image heterogeneous graph. Then, after removing the central person from the person connection diagram, community discovery is performed to obtain the social circle. After obtaining the social circle, based on the identified central person and the person-image heterogeneous graph, the type of each social circle is determined. On the one hand, the above process can identify the social circle and the type of social circle based on the face clustering result without relying on user annotation, with high intelligence and improved user experience. On the other hand, this process can accurately identify the central person without obtaining information such as the user's face unlocking, and then identify the social circle and the type of social circle, without affecting the security of the user's private information, further improving the user experience. On the third hand, the above process can be completed offline based on the images in the image library database without relying on the network. Compared with identifying the social circle and the type of social circle through social behaviors on the social network, this method has stronger applicability, and the recognition result is more targeted for the user of the electronic device, thus further improving the user experience. On the fourth hand, the above process first identifies the social circle, then determines the type of social circle based on the recognition result of the social circle, and corrects the relationship type of the persons in the social circle. Compared with separately classifying the relationships between persons to obtain different social circles and the types of social circles, the method provided in this application has higher accuracy in identifying the social circle and the type of social circle and better user experience.

[0419] It should be noted that the order of the methods provided in the embodiments of the present application is only an example and does not impose any limitation on the visual media search method. In practical applications, when the implementation logic is met, there can be various orders between steps. For example, two steps can be executed one after the other or simultaneously. And when executed one after the other, when the implementation logic is met, the order between the two steps is not limited either. For example, in the above embodiment, for steps S402 to S404 and step S405, it is possible to first execute steps S402 to S404 and then execute step S405, or first execute step S405 and then execute steps S402 to S404, or execute step S405 while executing steps S402 to S404.

[0420] It can be understood that after determining the social circle and the type of the social circle in the above manner, the type of the social circle to which the incremental face image and the face representative frame belong can be queried according to the identification information of the incremental face image and the face representative frame. In this way, the relationship information of the characters in the incremental face image and the face representative frame can be determined. For example, according to the identification information of a certain face representative frame x, it is queried that the representative frame x belongs to the social circle M, and the type of the social circle M is colleagues. Therefore, the relationship information of the character of the representative frame x can be determined as colleagues, and further the relationship information of the characters in the video segment to which the representative frame x belongs can be determined as colleagues.

[0421] It should be noted that some incremental face images or face representative frames may not belong to any social circle, so the relationship information of the characters cannot be determined by the above method, that is, the relationship information of the characters can be empty. And some incremental face images or face representative frames can not only determine the relationship information of the characters by the above method, but also have the relationship information of the characters actively added by the user. In this case, the relationship information obtained by the two methods can be stored, providing a wider search range for subsequent searches and improving search accuracy. All in all, the relationship information of the characters in a certain image or video segment may be empty, may include one kind of relationship information, or may include multiple kinds of relationship information, which can be determined according to the actual situation.

[0422] The process of determining the intimacy between two characters will be described below.

[0423] In a specific embodiment, taking any two characters included in the images in the co-occurrence image set, namely character a and character b, as an example, the process of calculating the intimacy between character a and character b will be described. In the implementation process of the above step “S403. According to the attribute information of each image in the co-occurrence image set, determine the intimacy between pairwise characters in the face clustering result”, according to the attribute information of each image in the co-occurrence image set, determining the intimacy between pairwise characters included in the images in the co-occurrence image set includes:

[0424] 1) Calculate the mutual information between person a and person b.

[0425] Optionally, the mutual information PMI(a, b) between person a and person b can be calculated by formula (4):

[0426]

[0427] where pic a represents the number of co-occurrence images containing person a in the co-occurrence image set. pic b represents the number of co-occurrence images containing person b in the co-occurrence image set. pic a,b represents the number of co-occurrence images of person a and person b in the co-occurrence image set. total_pics represents the total number of all co-occurrence images in the co-occurrence image set.

[0428] In one embodiment, the intimacy between person a and person b can be directly characterized by PMI(a, b), that is, the value of PMI(a, b) is the intimacy between person a and person b.

[0429] In some other embodiments, the intimacy between person a and person b can also be determined based on PMI(a, b) in combination with other parameters. As a possible implementation, based on PMI(a, b), in combination with the time distribution divergence and location distribution divergence of the images, the intimacy between person a and person b can be determined according to the following process. Specifically, on the basis of the above step 1), the following steps 2) to 6) can also be included:

[0430] 2) According to the acquisition time of the co-occurrence images of person a and person b, count the time distribution of the co-occurrence images of person a and person b (hereinafter simply referred to as the co-occurrence time distribution of person a and person b).

[0431] The co-occurrence time distribution of person a and person b is used to characterize the distribution of the acquisition times of the co-occurrence images of person a and person b in multiple preset time intervals. The interval length and the number of intervals of the preset time intervals can be determined according to the actual situation.

[0432] As a possible implementation, based on the timing information, a day (00:00 to 24:00) can be divided into multiple preset time intervals, where the duration of each preset time interval is less than 24 hours. In a specific embodiment, a day can be divided into 24 preset time intervals, and the duration of each preset time interval is 1 hour. For example, the 24 preset time intervals can include: [00:00, 01:00), [01:00, 02:00), [02:00, 03:00)... [23:00, 24:00).

[0433] It should be noted that the preset time interval can also be divided in other ways. For example, based on the date information and the timing information, a month (from 00:00 on the 1st to 24:00 on the 31st) can be divided into multiple preset time intervals. The present application does not make any limitation on the division method of the preset time interval. Hereinafter, taking the division of one day into multiple preset times based on the timing information as an example for illustration, the execution process of the method under other division methods is similar and will not be elaborated.

[0434] In a specific embodiment, the co-occurrence time distribution of characters a and b can be counted according to the following process:

[0435] a. Count the number of matching co-occurrence images in each preset time interval in the co-occurrence images of characters a and b.

[0436] Among them, the matching co-occurrence image in a certain preset time interval refers to: in the co-occurrence image of characters a and b, the image whose timing information belongs to this preset time interval (that is, the acquisition moment matches this preset time interval). For example, the matching image in the preset time interval [08:00, 09:00) refers to the co-occurrence image in the co-occurrence image of characters a and b whose timing information belongs to [08:00, 09:00). For example, the acquisition moment of a co-occurrence image x of characters a and b (abbreviated as image x) is 08:18:12 on April 12, 2023, and the timing information of image x is 08:18:12. Since this timing information belongs to [08:00, 09:00), this image x is a matching co-occurrence image in the preset time interval [08:00, 09:00). According to this method, the matching co-occurrence images in each preset time interval are determined, and the number of matching co-occurrence images in each preset time interval is counted.

[0437] Optionally, during the statistical process, for multiple matching co-occurrence images with the same date information within the same preset time interval, the quantity may be counted only once. That is: the person ab co-occurrence images with the same date information and the timing information belonging to the same preset time interval are counted as one image. For example, the acquisition time of image x is 08:18:12 seconds on April 12, 2023, and the acquisition time of the person ab co-occurrence image y (referred to as image y) is 08:25:10 seconds on April 12, 2023. The date information of image x and image y is the same, both being April 12, 2023. The timing information of image x and image y both belongs to the preset time interval [08:00, 09:00), and they are matching co-occurrence images within this preset time interval. Therefore, during the statistics, image x and image y are regarded as one image, and the quantity is counted only once. This is because the co-occurrence time distribution of person ab will be used later to determine the divergence (i.e., the degree of uniformity) of the time distribution of person ab, and the intimacy between person a and person b is reflected through the time distribution divergence. The more uniform the time distribution, the closer person a and person b are. In actual use, even for two very close people, when taking pictures with a camera, they may concentrate on taking pictures within a certain preset time interval on a certain day, resulting in multiple images. And adding the concentrated pictures taken within a certain preset time interval to the time distribution statistics will cause the co-occurrence time distribution of the two people to be uneven, which may make the intimacy of two relatively close people lower, while the intimacy of two people who do not have concentrated pictures taken and are not close may increase instead. Therefore, through the method provided by this implementation method, the pictures taken during concentrated shooting are counted only once, and the pictures generated by concentrated shooting are excluded to prevent affecting the accuracy of the co-occurrence time distribution statistics of person ab, and thus the accuracy of intimacy calculation can be improved.

[0438] In a specific embodiment, taking multiple preset time intervals including: [00:00, 01:00), [01:00, 02:00), [02:00, 03:00) …… [23:00, 24:00) as an example, the process of counting the quantity of matching co-occurrence images within each preset time interval can be expressed by the following formula:

[0439] Through the vector represents the quantity of matching co-occurrence images in each preset time interval. t ab [i] represents the element in the vector t ab where i represents the preset time interval serial number, i = 1, 2, 3 … 24. Among them, when i = 1, the preset time interval 1 can be [00:00, 01:00), when i = 2, the preset time interval 2 can be [01:00, 02:00), when i = 3, the preset time interval 3 can be [02:00, 03:00) …… when i = 24, the preset time interval 24 can be [23:00, 24:00). t ab[i] represents the number of matching images within the preset time interval i.

[0440] First, initialize each element in the vector t ab to 0. Then, based on formula (5), count the number of matching co-occurring images within each preset time interval:

[0441]

[0442] where m represents the co-occurring image m of person ab (abbreviated as image m). The function f(m) represents the preset time interval to which the timing information of image m belongs. t ab [f(m)] represents the number of matching images within the preset time interval f(m). g(m) = 1 indicates that the co-occurring image of person ab with the same date information as image m and belonging to the same preset time interval has been counted once. In this case, it is no longer counted. Therefore, the value of t ab [f(m)] remains unchanged, that is, t ab [f(m)] = t ab [f(m)]. Correspondingly, g(m) = 0 indicates that the co-occurring image of the person with the same date information as image m and belonging to the same preset time interval has not been counted. In this case, this image needs to be included in the statistics. Therefore, the value of t ab [f(m)] is incremented by 1, that is, t ab [f(m)] = t ab [f(m)] + 1.

[0443] M represents the set of co-occurring images of person ab, M1, M2... Mx all represent the co-occurring images of person ab, and x represents the total number of co-occurring images of person ab.

[0444] It can be understood that in the case where multiple preset time intervals include: [00:00, 01:00), [01:00, 02:00), [02:00, 03:00)... [23:00, 24:00), when determining which preset time interval the timing information of a certain image belongs to, it can be judged only by the hour information of the image without obtaining the minute information and second information, etc. For example, if the hour information of image m is 10, then image m belongs to the preset time interval [10:00, 11:00). Therefore, in this embodiment, multiple preset time intervals are divided into 24 preset time intervals in the above manner, and the lower limit of each preset time interval is an integer hour, which can reduce the calculation amount during time distribution statistics and improve the algorithm operation efficiency.

[0445] b. Determine the co-occurrence time distribution of person ab according to the number of matching co-occurring images within each preset time interval.

[0446] Based on the number of matching co-occurrence images in each determined preset time interval, determine the proportion of the matching co-occurrence images in all person ab co-occurrence images in each preset time interval, and obtain the co-occurrence time distribution of person ab.

[0447] For the sake of convenience of description, hereinafter, the proportion of the number of matching co-occurrence images in a certain preset time interval in all person ab co-occurrence images is simply referred to as the co-occurrence image proportion corresponding to this preset time interval.

[0448] Optionally, the process of determining the co-occurrence time distribution of person ab can be expressed by formula (6):

[0449]

[0450]

[0451] Among them, the vector represents the co-occurrence time distribution of person ab, and d ab,t [i] is any element in the vector, and d ab,t [i] represents the co-occurrence image proportion corresponding to the preset time interval i.

[0452] As an example, the co-occurrence time distribution diagram of person ab can be as Figure 21 shown. Among them, the vertical coordinate represents the co-occurrence image proportion corresponding to the preset time interval (simply referred to as proportion in the figure), and the horizontal coordinate represents the serial number of the preset time interval, that is, the value of i. Figure 21 The co-occurrence time distribution of person ab shown in

[0453] 3) Determine the KL divergence between the co-occurrence time distribution of person ab and the uniform time distribution to obtain the time KL divergence of person ab.

[0454] The uniform time distribution means that all person ab co-occurrence images are evenly distributed in multiple preset time intervals, that is, the co-occurrence image proportion corresponding to each preset time interval is the same, and is all 1 / the number of preset time intervals. For example, when the number of preset time intervals is 24, the uniform time distribution means that the co-occurrence image proportion corresponding to each preset time interval is all 1 / 24.

[0455] The time KL divergence of person ab is used to characterize the difference between the co-occurrence time distribution of person ab and the uniform time distribution. In other words, the time KL divergence of person ab is used to characterize the degree of uniformity of the distribution of person ab co-occurrence images in multiple preset time intervals, so as to be able to reflect the degree of uniformity of the co-occurrence of person a and person b in time.

[0456] The time uniform distribution can be expressed as a vector The elements in this set are all 1 / 24. Then, the temporal KL divergence of person ab can be calculated by formula (7):

[0457]

[0458] Among them, D ab,t represents the temporal KL divergence of person ab, and d u,t [i] represents any element in the vector d u,t , and d u,t [i] = 1 / 24.

[0459] As an example, a comparison schematic diagram of the co-occurrence time distribution and the uniform time distribution of person ab can be as Figure 22 shown. The meanings of the abscissa and ordinate in the figure are the same as those in Figure 21 and will not be elaborated here.

[0460] 4) Statistically analyze the location distribution of the co-occurrence images of person ab (hereinafter simply referred to as the co-occurrence location distribution of person ab).

[0461] The co-occurrence location distribution of person ab is used to characterize the distribution of the acquisition locations of the co-occurrence images of person ab among multiple known locations. Among them, the known locations can be obtained from the acquisition locations of all visual media in the CV database or the gallery database. It should be noted that "all visual media" includes both visual media containing people and visual media not containing people.

[0462] In a specific embodiment, taking the known locations from the CV database as an example for illustration. As in the above step S102, the gallery APP synchronizes the incremental visual media and their attributes to the CV database each time. Therefore, all visual media in the gallery database can be stored in the CV data. Therefore, the co-occurrence location distribution of person ab can be statistically analyzed based on the CV database according to the following process:

[0463] a. Determine the location set, which includes the acquisition locations of all visual media in the CV database.

[0464] Optionally, the location set can be expressed as is an element of set N, representing the acquisition location. For example, the location set y represents the total number of acquisition locations.

[0465] Optionally, at any time before performing step a, the location set can be generated by searching for the acquisition locations of all visual media stored in the CV database and stored in the CV database. And when the visual media in the CV database is updated, the location set can be updated to facilitate obtaining the location set as needed when implementing the method provided in this application embodiment.

[0466] For ease of description, the elements in the location set are hereinafter simply referred to as location elements. That is to say, the location set includes multiple location elements, and each location element is a location where a visual medium in the CV database is obtained.

[0467] b. Count the number of co-occurrence images of person ab corresponding to each location element.

[0468] For any location element Ni, the number of co-occurrence images of person ab corresponding to this location element refers to the number of co-occurrence images of person ab whose acquisition location is Ni. For example, the number of co-occurrence images of person ab corresponding to the location element "Beijing" refers to the number of co-occurrence images of person ab whose acquisition location is "Beijing".

[0469] c. Determine the co-occurrence location distribution of person ab according to the number of co-occurrence images of person ab corresponding to each location element.

[0470] Based on the number of co-occurrence images of person ab corresponding to each location element determined above, determine the proportion of the number of co-occurrence images of person ab corresponding to each location element in all co-occurrence images of person ab, and obtain the co-occurrence location distribution of person ab.

[0471] For ease of description, hereinafter, the proportion of the number of co-occurrence images of person ab corresponding to a certain location element in all co-occurrence images of person ab is simply referred to as the co-occurrence image proportion corresponding to this location element.

[0472] Optionally, when determining the co-occurrence image proportion corresponding to each location element, if the number of co-occurrence images of person ab corresponding to the location element is 0, then the proportion is determined to be 0; if the number of co-occurrence images of person ab corresponding to the location element is not 0 (i.e., greater than 0), then the co-occurrence image proportion corresponding to this location element is determined to be 1 / the number of non-zero location elements. Among them, the non-zero location elements refer to the location elements whose corresponding number of co-occurrence images of person ab is not 0. For example, the location set includes 10 location elements. After statistics, the number of co-occurrence images of person ab corresponding to 5 location elements N1, N3, N5, N7, and N10 in this set is not 0, then the co-occurrence image proportion corresponding to these 5 location elements is all 1 / 5 = 0.2, and the co-occurrence image proportion corresponding to other location elements is 0. Then, the co-occurrence location distribution of person ab can be expressed as:

[0473] As another alternative implementation manner, in the process of counting the location distribution of co-occurrence images of person ab, steps b and c above can also be replaced by: counting whether there are co-occurrence images of person ab corresponding to each location element; determining the co-occurrence location distribution of person ab according to the statistical results.

[0474] Optionally, the tag set includes tag values corresponding to each location element, and the tag value is used to mark whether there is an image with the acquisition location being the location element in the set of co-occurrence images of person a and person b. Specifically, for any location element Ni, it is determined whether there is an image with the acquisition location being Ni in the set of co-occurrence images of person a and person b; if so, the tag value corresponding to the location element Ni is determined to be 1 (i.e., the first value); if not, the tag value corresponding to the location element Ni is determined to be 0 (i.e., the second value). Then, the total number of location elements with a tag value of 1 is counted to obtain the number of non-zero location elements (i.e., the total number of the first values). Then, the proportion of co-occurrence images corresponding to the location elements with a tag value of 1 is set to 1 / the number of non-zero location elements.

[0475] Specifically, the above process can be expressed by the following formula:

[0476] Through the vector represents the set of tag values corresponding to each location element. l ab [j] represents the element in the vector l ab , the serial number of the j location element, j = 1, 2, 3... y. l ab [j] represents the tag value corresponding to the location element j.

[0477] First, each element in the vector l ab is initialized to 0. Then, based on formula (8), values are assigned to each element in the vector l ab :

[0478]

[0479] Among them, m represents the co-occurrence image m of person a and person b (abbreviated as image m). The function h(m) represents the acquisition location of image m. l ab [h(m)] represents the tag value corresponding to the acquisition location h(m). q(m) = 0 means that the acquisition location h(m) has not been counted, and the tag value of the location element with the same acquisition location h(m) is currently 0. In this case, since the acquisition location of image m is h(m), therefore, the tag value corresponding to the location element with the same acquisition location h(m) is assigned the value 1, that is, l ab [h(m)] = 1. Correspondingly, q(m) = 1 means that the acquisition location h(m) has been counted once, that is, it has been determined that there is a co-occurrence image of person a and person b with the acquisition location being h(m), and the tag value of the location element with the same acquisition location h(m) has been assigned the value 1. In this case, l ab [h(m)] remains unchanged, that is, l ab [h(m)] = l ab [h(m)]. M represents the set of co-occurrence images of person a and person b, which will not be elaborated here.

[0480] After that, the co-occurrence location distribution of person ab is determined by formula (9):

[0481]

[0482]

[0483] where the vector represents the co-occurrence location distribution of person ab, and d ab,l [j] is any element in the vector, and d ab,l [j] represents the proportion of co-occurrence images corresponding to location element j.

[0484] In this implementation, considering that the co-occurrence location distribution of person ab will be used later to determine the divergence (i.e., uniformity) of the location distribution of person ab, the intimacy between person a and person b is reflected by the location distribution divergence. The more uniform the location distribution is, the closer person a and person b are. In actual use, even for two very intimate people, when taking pictures with a camera, they may concentrate on taking pictures at a certain location, resulting in multiple images. And adding the concentrated photography at a certain location to the location distribution statistics will lead to the non-uniformity of the co-occurrence location distribution of the two people, which may make the intimacy of two relatively close people lower, while the intimacy of two people who are not close and do not have concentrated photography may increase instead. Therefore, in this implementation, by counting whether there are co-occurrence images of person ab corresponding to each location element, the location distribution of person ab is analyzed without considering the number of images obtained at each location element, preventing the images generated by concentrated photography from affecting the accuracy of the co-occurrence location distribution statistics of person ab, and thus improving the accuracy of intimacy calculation.

[0485] As an example, when the co-occurrence location distribution of person ab is , the corresponding co-occurrence location distribution diagram can be as Figure 23 shown. Among them, the vertical coordinate represents the proportion of co-occurrence images corresponding to the location element, and the horizontal coordinate represents the serial number of the location element, that is, the value of j.

[0486] 5) Determine the KL divergence between the co-occurrence location distribution of person ab and the uniform location distribution to obtain the location KL divergence of person ab.

[0487] The uniform location distribution means that all co-occurrence images of person ab are evenly distributed among all location elements in the location set, that is, the proportion of co-occurrence images corresponding to each location element is the same, and is 1 / total number of location elements. For example, when the total number of location elements is 10, the uniform location distribution means that the proportion of co-occurrence images corresponding to each location element is 1 / 10 = 0.1.

[0488] The Kullback-Leibler divergence of the locations of person ab is used to characterize the difference between the co-occurrence location distribution of person ab and the uniform location distribution. In other words, the Kullback-Leibler divergence of the locations of person ab is used to characterize the degree of uniformity of the distribution of co-occurrence images of person ab in multiple preset time intervals, thereby being able to reflect the degree of uniformity of the co-occurrence of person a and person b in terms of location.

[0489] The location uniform distribution can be represented as a vector Each element in this set is 1 / y. Then, the Kullback-Leibler divergence of the locations of person ab can be calculated by formula (10):

[0490]

[0491] where D ab,l represents the Kullback-Leibler divergence of the locations of person ab, and d u,l [j] represents any element in the vector d u,l , and d u,l [j] = 1 / y.

[0492] As an example, a comparison schematic diagram of the co-occurrence location distribution and the uniform location distribution of person ab can be as Figure 24 shown. The meanings of the abscissa and ordinate in the figure are the same as Figure 23 and will not be elaborated here.

[0493] 6) Determine the intimacy between person a and person b based on the temporal Kullback-Leibler divergence of person ab, the location Kullback-Leibler divergence of person ab, and PMI(a, b).

[0494] In a specific embodiment, the intimacy between person a and person b can be determined according to formula (11):

[0495] closure(a, b) = PMI(a, b) × (1 - D ab,t ) × (1 - D ab,l ) (11)

[0496] where closure(a, b) represents the intimacy between person a and person b, D ab,t represents the temporal Kullback-Leibler divergence of person ab, and D ab,l represents the location Kullback-Leibler divergence of person ab.

[0497] It can be seen from formula (11) that the smaller D ab,t , the larger closure(a, b); the smaller D ab,l , the larger closure(a, b). The smaller D ab,t means that the co-occurrence time distribution of person ab is more uniform; the smaller D ab,lThe smaller it is, the more evenly distributed the co-occurrence locations of person ab are. That is to say, according to formula (11), the more evenly distributed the co-occurrence time of person ab is, the greater the intimacy between person a and person b, and the closer person a and person b are; moreover, the more evenly distributed the co-occurrence time and location of person ab are, the greater the intimacy between person a and person b, and the closer person a and person b are.

[0498] In this embodiment, when determining the intimacy between person a and person b, on the basis of PMI(a,b), the KL divergence of the co-occurrence time and the KL divergence of the co-occurrence location of person ab are further added. The KL divergence of the co-occurrence time of person ab can characterize the degree of evenness of the distribution of the co-occurrence images of person a and person b in time. The KL divergence of the co-occurrence location of person ab can characterize the degree of evenness of the distribution of the co-occurrence images of person a and person b in location. The closer the social relationship between person a and person b is, the more evenly distributed their co-occurrence images are in terms of time and location. Therefore, in this embodiment, the KL divergence of the co-occurrence time and the KL divergence of the co-occurrence location of person ab are fused into the calculation of intimacy, improving the accuracy of intimacy calculation, and thus improving the accuracy of social circle recognition.

[0499] It can be understood that the above steps 2) to 5) are also called divergence calculation operations, and the above steps 1) to 6) are also called intimacy calculation operations.

[0500] The method for identifying the central person will be described below.

[0501] In a specific embodiment, in the above step S405, according to the identified person-heterogeneous graph, identifying the central person may specifically include:

[0502] 1) According to the person-image heterogeneous graph, determine the image set of each person.

[0503] As described in the above embodiment, an image containing a certain person can be called the image of that person.

[0504] Specifically, according to the person-image heterogeneous graph, all image nodes connected to a certain person node can be determined, and a set of images corresponding to these image nodes can be generated, that is, the image combination of the person corresponding to the person node is obtained. Taking any person a as an example, according to the connection relationship matrix shown in Table 5, the connection matrix can be determined, and the image nodes including person node a in the connection matrix can be determined: image node A, image node C, and image node N, etc. The images corresponding to these image nodes are generated into a set to obtain the image set of person a.

[0505] 2) According to the attribute information of each image in the image set of person a, determine the KL divergence of the time of the image of person a (hereinafter simply referred to as the time KL divergence of person a) and the KL divergence of the location of the image of person a (hereinafter simply referred to as the location KL divergence of person a).

[0506] Among them, person a is any person in the person-image heterogeneous graph. That is to say, for each person in the person-image heterogeneous graph, the temporal KL divergence and the spatial KL divergence of each person are determined according to the method of this step, so as to obtain the temporal KL divergence and the spatial KL divergence of each person.

[0507] In a specific embodiment, the temporal KL divergence and the spatial KL divergence of person a can be determined according to the following process:

[0508] a. According to the acquisition times of the images in the image set of person a, count the temporal distribution of the images of person a (hereinafter referred to as the temporal distribution of person a).

[0509] The temporal distribution of person a is used to characterize the distribution of the acquisition times of the images of person a in multiple preset time intervals. The preset time intervals are similar to those in the above embodiment and will not be elaborated here.

[0510] Counting the temporal distribution of person a is similar to counting the co-occurrence temporal distribution of person ab in the above embodiment. The difference is that: in this embodiment, the statistical range is the image set of person a, while in the above embodiment, the statistical range is all the person ab co-occurrence images; in this embodiment, the statistical object is the images of person a, while in the above embodiment, the statistical object is the person ab co-occurrence images. In this embodiment, among the image set of person a, the images whose acquisition times match a certain preset time interval can be called the matching images of this preset time interval. While in the above embodiment, the images whose acquisition times match a certain preset time interval can be called the matching co-occurrence images of this preset time interval. For the specific implementation process and beneficial effects of this embodiment, please refer to the above embodiment and will not be elaborated here.

[0511] b. Determine the KL divergence between the temporal distribution of person a and the uniform temporal distribution to obtain the temporal KL divergence of person a.

[0512] The uniform temporal distribution is the same as that in the above embodiment. The temporal KL divergence of person a is used to characterize the difference between the temporal distribution of person a and the uniform temporal distribution. In other words, the temporal KL divergence of person a is used to characterize the degree of uniformity of the distribution of the images of person a in multiple preset time intervals, that is, to reflect the degree of uniformity of the appearance of person a in time.

[0513] The calculation method of the temporal KL divergence of person a is similar to that of the temporal KL divergence of person ab and will not be elaborated here.

[0514] c. According to the acquisition locations of the images in the image set of person a, count the spatial distribution of the images of person a (hereinafter referred to as the spatial distribution of person a).

[0515] The location distribution of person a is used to characterize the distribution of the acquisition locations of the images of person a.

[0516] In a specific embodiment, the co-occurrence location distribution of persons a and b can be counted according to the following process:

[0517] Counting the location distribution of person a is similar to counting the co-occurrence location distribution of persons a and b in the above embodiment. The difference is that in this embodiment, the statistical scope is the image set of person a, rather than the co-occurrence image set of persons a and b, and the statistical object is the image of person a rather than the co-occurrence image of persons a and b. For the specific implementation process and beneficial effects of this process, please refer to the above embodiment and will not be elaborated here.

[0518] d. Determine the KL divergence between the location distribution of person a and the uniform location distribution to obtain the location KL divergence of person a.

[0519] The uniform location distribution is the same as that in the above embodiment. The location KL divergence of person a is used to characterize the difference between the location distribution of person a and the uniform location distribution. In other words, the location KL divergence of person a is used to characterize the degree of uniformity of the distribution of the images of person a in multiple preset time intervals, so as to reflect the degree of uniformity of the appearance of person a in terms of location.

[0520] The calculation method of the location KL divergence of person a is similar to that of the location KL divergence of persons a and b, and will not be elaborated here.

[0521] 3) Determine the central person according to the time KL divergence and location KL divergence of all persons.

[0522] In an embodiment, the time KL divergence and location KL divergence of each person can be summed to obtain the divergence sum of each person. Then, rank the divergence sums of all persons (i.e., the first ranking), and take the Q1 persons with the smallest divergence sum as the central persons, where Q1 is an integer greater than or equal to 1. It can be understood that when summing the KL divergences, it can be a direct sum or a weighted sum, and the embodiments of the present application do not make any limitations on this.

[0523] In another embodiment, the KL divergence of time for all persons can be ranked in ascending order, and the KL divergence of location can be ranked in ascending order. Then, the ranking of the KL divergence of time and the ranking of the KL divergence of location for each person are summed to obtain the sum of divergence rankings corresponding to each person. The Q2 persons with the smallest sum of all divergence rankings are taken as the central persons, where Q2 is an integer greater than or equal to 1. For example, the rankings of the KL divergence of time for persons A, B, and C are {1, 2, 3}, and the rankings of the KL divergence of location are {2, 3, 1}. Then, the two rankings are summed to obtain the sum of divergence rankings for persons A, B, and C as: {3, 5, 4}. If Q2 is 1, then person A with the smallest sum of divergence rankings is taken as the central person. It should be noted that when determining the central persons, if it is impossible to directly obtain Q2 central persons due to the same sum of divergence rankings for multiple persons, then persons with a smaller KL divergence of time and / or KL divergence of location are selected from the multiple persons as the central persons. In addition, when summing the rankings, it can be a direct sum or a weighted sum, and the embodiments of the present application do not make any limitations in this regard.

[0524] That is to say, persons whose KL divergence of time and KL divergence of location meet the preset conditions are determined as the central persons, where the preset conditions are any one of the following conditions: ① The sum of divergences ranks among the top Q1 in the first ranking, and the first ranking is obtained by ranking persons in ascending order of the sum of divergences; ② The sum of divergence rankings ranks among the top Q2 in the second ranking, and the second ranking is obtained by ranking persons in ascending order of the sum of divergence rankings; ③ The sum of divergences ranks among the last Q1 in the fourth ranking, and the fourth ranking is obtained by ranking persons in descending order of the sum of divergences; ④ The sum of divergence rankings ranks among the last Q2 in the fifth ranking, and the fifth ranking is obtained by ranking persons in descending order of the sum of divergence rankings.

[0525] In another embodiment of the present application, collection recommendations can also be made using the social circles identified by the above method. For example, according to the persons included in a certain social circle (referred to as target persons), visual media including the target persons in the gallery database can be selected to generate a collection, and according to the type of the social circle, the theme of the collection can be generated. Then, the collection is recommended to the user. For example, visual media including the persons in social circle 1 are selected to generate collection 1. The type of social circle 1 is family, so the theme of collection 1 is set as "a family". Exemplarily, the interface including the "a family" collection can be as Figure 25 shown.

[0526] Optionally, when generating a collection based on a social circle, it can further be generated in combination with the acquisition time information of visual media, the acquisition location of visual media, etc. Moreover, when generating the theme of the collection, it can also be generated in combination with the acquisition time information of visual media and the acquisition location of visual media. For example, generate a collection with the theme of "Memories with Family during the Spring Festival" from the visual media in social circle 1 whose acquisition time is within the preset Spring Festival period.

[0527] The above text details an example of the visual media search method provided by the embodiments of this application. It can be understood that for an electronic device to implement the above functions, it includes the corresponding hardware and / or software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed herein, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of this application.

[0528] The embodiments of this application can divide the functional modules of the electronic device according to the above method examples. For example, each function can be corresponding to each functional module, such as a detection unit, a processing unit, a display unit, etc., or two or more functions can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of this application is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0529] It should be noted that all relevant contents of each step involved in the above method embodiments can be cited in the functional descriptions of the corresponding functional modules, and will not be repeated here.

[0530] The electronic device provided in this embodiment is used to execute the above visual media search method, so it can achieve the same effect as the above implementation method.

[0531] In the case of adopting an integrated unit, the electronic device can also include a processing module, a storage module, and a communication module. Among them, the processing module can be used to control and manage the actions of the electronic device. The storage module can be used to support the electronic device to execute stored program codes and data, etc. The communication module can be used to support the communication of the electronic device with other devices.

[0532] Among them, the processing module can be a processor or a controller. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, and so on. The storage module can be a memory. The communication module can specifically be a device that interacts with other electronic devices, such as a radio frequency circuit, a Bluetooth chip, a Wi-Fi chip, etc.

[0533] In one embodiment, when the processing module is a processor and the storage module is a memory, the electronic device involved in this embodiment can be a device having Figure 2 the structure shown.

[0534] The embodiments of the present application also provide a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the processor is caused to execute the visual media search method of any of the above embodiments.

[0535] The embodiments of the present application also provide a computer program product. When the computer program product runs on a computer, the computer is caused to execute the above-related steps to implement the visual media search method in the above embodiments.

[0536] In addition, the embodiments of the present application also provide a device, which can specifically be a chip, a component, or a module. The device can include a processor and a memory connected thereto; among them, the memory is used to store computer execution instructions. When the device runs, the processor can execute the computer execution instructions stored in the memory, so that the chip executes the visual media search method in each of the above method embodiments.

[0537] Among them, the electronic device, computer-readable storage medium, computer program product, or chip provided in this embodiment is all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.

[0538] Through the description of the above embodiments, those skilled in the art can understand that, for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0539] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0540] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0541] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0542] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks or optical discs and other various media that can store program codes.

[0543] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for visual media search, which is executed by an electronic device, characterized in that, the method includes: displaying a first interface, which includes a search box; receiving a first operation in which a user enters search text in the search box; the search text includes first information characterizing a person relationship; in response to the first operation, displaying a first search result; the first search result corresponds to a first video, the first video includes a first person, and the relationship between the first person and the central person is consistent with the first information, and the central person refers to the person in the visual media stored in the electronic device who is in the social center.

2. The method according to claim 1, characterized in that, the method further includes: in response to the first operation, displaying a second search result, the second search result corresponds to a first image; the first image includes a second person, and the relationship between the second person and the central person is consistent with the first information.

3. The method according to claim 1 or 2, characterized in that, the first video includes a plurality of second video segments, and at least one hit video segment is included in the plurality of second video segments, and the hit video segment is a video segment including the first person among the plurality of second video segments.

4. The method according to claim 3, characterized in that, the method further includes: receiving a second operation in which a user clicks the first search result; in response to the second operation, playing the first video starting from the first frame image of the first hit video segment, or playing the first video starting from the representative frame image; the first hit video segment is one of the at least one hit video segment, and the representative frame image is a frame image in the first hit video segment, and the representative frame image includes the first person.

5. The method according to claim 4, characterized in that, the displaying the first search result includes: displaying a thumbnail of the cover frame image; the cover frame image is a frame image in the first video, and a first playback time is displayed in the thumbnail, and the first playback time is within the closed interval formed by the start playback time and the end playback time of the first hit video segment.

6. The method according to claim 5, characterized in that, the first playback time is: the start playback time of the first hit video segment, the end playback time of the first hit video segment, the middle playback time of the first hit video segment, or the playback time corresponding to the representative frame image.

7. The method according to claim 5 or 6, characterized in that, the cover frame image is the first frame image of the first video, or the first frame image of the first hit video segment, or the representative frame image.

8. A method for visual media search, which is executed by an electronic device, characterized in that, the method includes: displaying a first interface, which includes a search box; receiving a first operation in which a user enters search text in the search box; in response to the first operation, identifying first information characterizing a person relationship in the search text; Search for a first video among multiple videos according to the first information; the first video includes multiple second video segments, and at least one hit video segment is included in the multiple second video segments, and the person relationship information corresponding to the hit video segment is consistent with the first information; the person relationship information is used to characterize the relationship between the persons included in the visual media and the central person, and the central person refers to the person in the visual media stored in the electronic device who is in the social center; Display a first search result corresponding to the first video.

9. The method according to claim 8, wherein, the method further includes: Search for a first image among multiple images according to the first information; the person relationship information corresponding to the first image is consistent with the first information; Display a second search result corresponding to the first image.

10. The method according to claim 8 or 9, wherein, the method further includes: Perform natural language understanding processing on the search text to identify a first semantic subject and category information in the search text; Perform text semantic understanding on the search text to obtain a text semantic vector; The step of searching for a first video among multiple videos according to the first information includes: Perform multi-way recall in the index library according to the first information, the first semantic subject, the category information, and the text semantic vector to obtain a recall result, and the recall result includes the first video; the index library includes indexes of attribute information, person relationship information, and visual semantic vectors of multiple video segments and multiple images, and the attribute information includes at least one of the visual media acquisition time, the visual media acquisition location, the visual media classification label, and the visual media semantic subject.

11. The method according to claim 10, wherein, the method further includes: Receive a third operation for the user to add and / or modify the target visual media; the target visual media includes a target video and a target image; In response to the third operation, perform segmentation processing on the target video to obtain multiple third video segments; Extract a frame image from each of the third video segments as a representative frame image; Respectively obtain the person relationship information and visual semantic vectors of each of the target images and each of the representative frame images; Respectively obtain the attribute information of each of the target images and the attribute information of each of the target videos; Based on the person relationship information and visual semantic vectors of each of the target images and each of the representative frame images, and the attribute information of each of the target images and each of the target videos, construct indexes of the attribute information, person relationship information, and visual semantic vectors of each of the target images and each of the third video segments in the index library.

12. The method according to claim 11, wherein, The step of respectively obtaining the person relationship information and visual semantic vectors of each of the target images and each of the representative frame images includes: Perform a person relationship analysis on each of the target images and each of the representative frame images respectively to obtain the person relationship information corresponding to each of the target images and the person relationship information corresponding to each of the representative frame images; Perform an image semantic understanding on each of the target images and each of the representative frame images respectively to obtain the visual semantic vectors of each of the target images and the visual semantic vectors of each of the representative frame images.

13. The method according to claim 11 or 12, wherein, The construction of the indexes of the attribute information, person relationship information, and visual semantic vectors of each of the target images and each of the third video segments in the index library based on the person relationship information and visual semantic vectors of each of the target images and each of the representative frame images, and the attribute information of each of the target images and each of the target videos includes: Use the person relationship information corresponding to the first representative frame image as the person relationship information corresponding to the fourth video segment, where the first representative frame image is any one of the representative frame images, and the fourth video segment is the video segment to which the first representative frame image belongs; Use the visual semantic vector of the first representative frame image as the visual semantic vector of the fourth video segment; Based on the person relationship information and visual semantic vectors of each of the target images and each of the third video segments, and the attribute information of each of the target images and each of the target videos, construct indexes of the attribute information, person relationship information, and visual semantic vectors of each of the target images and each of the third video segments in the index library.

14. The method according to claim 12, wherein, The performing an image semantic understanding on each of the target images and each of the representative frame images respectively to obtain the visual semantic vectors of each of the target images and the visual semantic vectors of each of the representative frame images includes: Search for the person relationship information from the annotation information of the target image and the target video by the user to obtain the annotation relationship information; Identify the social circle and the type of the social circle according to the face image set; the face image set includes multiple face images, the face image is an image containing a face, the multiple face images include a target face image and a face representative frame, the target face image is an image containing a face in the target image, the face representative frame is an image containing a face in the representative frame, the social circle is a set of people having the same social relationship with the central person, and the type of the social circle represents the type of the social relationship between the people in the social circle and the central person; Determine the person relationship information corresponding to each of the target images and the person relationship information corresponding to each of the representative frame images according to the annotation relationship information, the social circle, and the type of the social circle.

15. The method according to claim 14, wherein, The identifying the social circle and the type of the social circle according to the face image set includes: Perform face clustering on the multiple face images to generate a face clustering result; the face clustering result characterizes the corresponding relationships among the multiple face images, multiple faces, and multiple persons. Determine a heterogeneous graph of person images according to the face clustering result; the heterogeneous graph of person images characterizes the corresponding relationships between the multiple face images and the multiple persons. Determine the intimacy between every two persons among the multiple persons according to the heterogeneous graph of person images, where the intimacy characterizes the degree of association between persons. Identify the social circle according to the intimacy between every two persons and the heterogeneous graph of person images. Determine the type of the social circle according to the heterogeneous graph of person images.

16. The method according to claim 15, wherein, the heterogeneous graph of person images includes multiple image nodes and multiple person nodes, the multiple image nodes correspond one-to-one to the multiple face images, and the multiple person nodes correspond one-to-one to the multiple persons; The identifying the social circle according to the intimacy between every two persons and the heterogeneous graph of person images includes: Removing the multiple image nodes in the heterogeneous graph of person images, and connecting the person nodes corresponding to two persons with non-zero intimacy by edges according to the intimacy between every two persons to obtain a person connection graph. Identify the central person according to the heterogeneous graph of person images. Removing the person node corresponding to the central person in the person connection graph to obtain a de-centered person connection graph. Perform community discovery on the de-centered person connection graph to obtain the social circle.

17. The method according to claim 15 or 16, wherein, The determining the type of the social circle according to the heterogeneous graph of person images includes: Inputting the heterogeneous graph of person images into a graph neural network model to extract the person feature vectors of each person among the multiple persons and the image feature vectors of each face image. Inputting the person feature vectors and the image feature vectors into a multi-layer perceptron model to predict the types of social relationships between each person among the multiple persons and the central person. Determine the number of persons corresponding to each type of social relationship in the social circle. Determine the type of social relationship with the largest corresponding number of persons as the type of the social circle.

18. The method according to any one of claims 9 to 17, wherein, In the visual media stored in the electronic device, the time distribution divergence and the location distribution divergence of the visual media containing the central person satisfy preset conditions, the time distribution divergence is the distribution divergence of the acquisition time of the visual media, and the location distribution divergence is the distribution divergence of the acquisition location of the visual media.

19. An electronic device, wherein, including: A processor, a memory, and an interface; The processor, the memory, and the interface cooperate with each other to enable the electronic device to execute the method according to any one of claims 1 to 18.

20. A computer-readable storage medium, wherein, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 18.

Citation Information

Cited By

  • Visual media search method, electronic device and storage medium

    WO2025103081A1