Retrieval method, readable storage medium, program product and electronic equipment
By employing a multimodal retrieval method that combines features from image, text, audio, and video data, the problem of low accuracy in single-modal retrieval is solved, resulting in more efficient image retrieval results and improved user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-07
AI Technical Summary
Existing image retrieval methods for electronic devices mainly rely on single-modal input, resulting in low retrieval accuracy. Users need to manually search for the required image from a large number of search results, which is time-consuming and inefficient.
A multimodal retrieval method is adopted, which detects retrieval data of multiple modalities (such as images, text, audio, and video), determines the target subject and its subject information of each modality, calculates the weight of each modality based on the amount of subject information and confidence level, and fuses their features to generate the first retrieval feature, thereby improving retrieval accuracy.
It improves the accuracy of image retrieval, reduces the time users spend searching for the desired image in the search results, and enhances the user experience.
Smart Images

Figure CN121808086A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of terminal, in particular to a retrieval method, readable storage medium, program product and electronic device. BACKGROUND
[0002] A large number of images can be stored in an electronic device (such as a mobile phone, a tablet, a computer, etc.). When a user shares an image or creates an image, the user needs to find a desired image from a large number of images. Therefore, the process of finding an image takes a long time, which brings a poor user experience to the user.
[0003] Currently, some electronic devices can support image retrieval, which can help a user quickly find a desired image. However, current image retrieval generally only supports input of a single modality (a source of information or a sense mode, such as audio, image, text, etc.). For example, only text retrieval of an image or only image retrieval of an image is supported. The accuracy of an image retrieved through a single modality is not high. Therefore, there are many images meeting a retrieval requirement in a retrieval result obtained through retrieval of a single modality. SUMMARY
[0004] Embodiments of the present application provide a retrieval method, readable storage medium, program product and electronic device to achieve the purpose of retrieving an image or a video based on multiple modalities and improve retrieval accuracy.
[0005] In a first aspect, embodiments of the present application provide a retrieval method applied to an electronic device, including: detecting a retrieval request including retrieval data of multiple modalities, determining at least part of subject information corresponding to a target subject in each retrieval data; performing feature fusion on a feature corresponding to each retrieval data based on a weight of each retrieval data, to obtain a first retrieval feature, wherein the weight of each retrieval data is related to a feature of at least part of the subject information corresponding to each retrieval data, and the feature of at least part of the subject information includes a quantity of at least part of the subject information; concentrating retrieval objects, and an object whose object feature matches the first retrieval feature is taken as a retrieval result of each retrieval data, and the greater the weight of the corresponding retrieval data is, the greater the relevance between the retrieval result and the retrieval data is.
[0006] In some embodiments of the present application, the retrieval request can include retrieval data of multiple modalities, such as image modalities, text modalities, audio modalities, video modalities, etc. The features corresponding to different retrieval data can be fused according to the weights of the retrieval data to obtain first retrieval features. The weights of the retrieval data are related to the features of the subject information provided by the retrieval data, and the features of the subject information can include the number of subject information. For example, the more the number of subject information provided, the greater the weight of the corresponding retrieval data. Therefore, when the features of each retrieval data are fused based on the weight, the retrieval data with a greater weight can dominate. Since the retrieval data with a greater weight provides more subject information, the first retrieval features after fusion can contain more subject information, and the retrieval results obtained based on the first retrieval features can better meet the retrieval requirements of the retrieval data, thereby improving the retrieval accuracy.
[0007] In a possible implementation of the first aspect, the modalities include at least one of the following modalities: image modalities, text modalities, audio modalities, and video modalities.
[0008] In a possible implementation of the first aspect, the subject information includes at least one of the following information: relative position relationship of target subjects, inter-object relationship of target subjects, biological action of target subjects, behavior of target subjects, traffic scene in which target subjects are located, motion scene of target subjects, relationship between target subjects and background, and number of target subjects.
[0009] In a possible implementation of the first aspect, the determining of the at least part of the subject information corresponding to the target subject in each retrieval data includes: obtaining a confidence of the subject information of the target subject in each retrieval data, the confidence being used to indicate the accuracy of the corresponding subject information; and regarding the subject information with a confidence greater than a confidence threshold in the subject information of each retrieval data as at least part of the subject information of the corresponding retrieval data.
[0010] In some embodiments of the present application, the confidence of the subject information corresponding to each retrieval data can be obtained, and then it is determined whether to retain the subject information according to the corresponding confidence threshold. In some embodiments, the accuracy threshold can be any value between 80% and 100%. It can be understood that since the subject information in each retrieval data can be relatively large, and some subject information is not accurate enough, it is necessary to screen the subject information corresponding to each retrieval data. It can be understood that the relationship between target subjects and actions described by the subject information greater than the confidence threshold are relatively accurate, and therefore the subject information greater than the confidence threshold can be regarded as at least part of the subject information of the retrieval data.
[0011] In one possible implementation of the first aspect above, determining at least some subject information corresponding to the target subject in each retrieval data further includes: converting each retrieval data into transformed retrieval data corresponding to the reference modality; determining the transformed subject information corresponding to the target subject in each transformed retrieval data; and using the transformed subject information as at least some subject information of the corresponding retrieval data.
[0012] In some embodiments of this application, retrieval data from multiple different modalities can be converted into retrieval data from the same reference modality, and the subject and subject information can be extracted separately to obtain the target subject and subject information corresponding to each modality.
[0013] For example, the text modality can be used as a reference modality. Image modality retrieval data can be converted into text describing the images, and audio modality retrieval data can be converted into corresponding text, thus enabling modal recognition of subject information within the text. It can be understood that identifying the target subject and subject information within the same modality has similar recognition benchmarks, making the weighting of each retrieval data point more accurate.
[0014] In one possible implementation of the first aspect above, the weight of each of the aforementioned search data is: the proportion of the number of at least part of the subject information corresponding to each search data to the number of at least part of the subject information of all search data.
[0015] In some embodiments of this application, the weight of each retrieval data is related to the characteristics of at least some of the subject information provided by each retrieval data.
[0016] For example, consider the amount of subject information. In a certain image retrieval process, three sets of search data are provided. The first set is in text mode, providing 3 pieces of subject information; the second set is in image mode, providing 4 pieces of subject information; and the third set is also in image mode, providing 3 pieces of subject information. In this case, the weight of the first query semantic could be 0.3, the weight of the second query semantic could be 0.4, and the weight of the third query semantic could be 0.3.
[0017] In other embodiments, the weight of each retrieval data can also be related to the amount of at least some subject information corresponding to the retrieval data and the confidence level of the subject information. For example, if there are two retrieval data, and the first retrieval data provides one subject information with a confidence level of 0.9, then 0.9 can be used as the contribution value for determining the weight of this retrieval data. The second retrieval data provides two subject information, with the first subject information having a confidence level of 0.8 and the second subject information having a confidence level of 0.9, then 0.8 + 0.9 = 1.7 can be used as the contribution value of this retrieval data. Therefore, the weight of the first retrieval data is 0.9 / (0.9 + 1.7) = 0.346, and the weight of the second retrieval data is 1.7 / (0.9 + 1.7) = 0.654.
[0018] In one possible implementation of the first aspect described above, the set of objects to be retrieved includes an image set and / or a video set.
[0019] It is understandable that when the retrieval object set is an image set, the similarity between the first retrieval feature and the image features of each image in the image set can be obtained, and then the image corresponding to the image feature that exceeds the similarity threshold can be selected as the retrieval result.
[0020] When the retrieval target set is a video set, at least one keyframe can be extracted from a video, and then the features of that keyframe can be obtained and used as the video features of that video. If multiple keyframes are extracted from a video, the features of the multiple keyframes can be fused to obtain the keyframes of the video. Then, the similarity between the first retrieval feature and the video features of each video in the video set can be obtained, and the video corresponding to the video feature that exceeds the similarity threshold can be selected as the retrieval result.
[0021] In one possible implementation of the first aspect above, determining at least some subject information corresponding to the target subject in each retrieval data includes: identifying at least one subject in the image of the retrieval data corresponding to the image modality; based on user operation, identifying at least some of the subjects in the at least one subject as the target subject corresponding to the retrieval data; and determining at least some subject information of the retrieval data of the image modality according to the target subject in the image.
[0022] In some embodiments of this application, since the image retrieval data may contain more subjects, the target subject can be determined from the image retrieval data after inputting the image retrieval data for the image modality. The target subject is the subject contained in the image or video that the user needs to retrieve.
[0023] For example, models for recognizing image subjects can identify various subjects in image retrieval data and provide them to the user for selection. Users can select a target subject or delete non-target subjects through voice, actions, gestures, clicks, etc., to determine the target subject in the image modality retrieval data. After determining the target subject provided in the image modality retrieval data, subject information for each target subject in the image can be obtained using an image encoder.
[0024] Secondly, this application provides an electronic device comprising: a memory for storing instructions; and at least one processor for executing the instructions to cause the device to implement the methods provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achievable in the second aspect can be referred to the beneficial effects of the methods provided in any embodiment of the first aspect, and will not be repeated here.
[0025] Thirdly, this application provides a computer-readable storage medium storing instructions that, when executed by a device, cause a computer to implement the methods provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achievable in this third aspect can be referenced to the beneficial effects of the methods provided in any embodiment of the first aspect, and will not be repeated here.
[0026] Fourthly, this application provides a computer program product that, when run on a device, enables the device to implement the methods provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achievable in the fourth aspect can be found in the beneficial effects of the methods provided in any embodiment of the first aspect, and will not be repeated here. Attached Figure Description
[0027] Figure 1 A schematic diagram illustrating the process of an electronic device retrieving images via text is shown.
[0028] Figure 2 According to some embodiments of this application, a schematic diagram of image retrieval via multimodal methods is shown;
[0029] Figure 3 According to some embodiments of this application, a flowchart of an implementation method for a retrieval method is shown;
[0030] Figure 4 According to some embodiments of this application, an image retrieval system is shown;
[0031] Figure 5A According to some embodiments of this application, a schematic diagram of selecting various target subjects is shown;
[0032] Figure 5BAccording to some embodiments of this application, a schematic diagram of deleting various non-target entities is shown;
[0033] Figure 6 According to some embodiments of this application, a schematic diagram of text and labels is shown;
[0034] Figure 7 According to some embodiments of this application, a schematic diagram of a target subject in audio extraction is shown;
[0035] Figure 8 According to some embodiments of this application, a schematic diagram of extracting keyframes from a video is shown;
[0036] Figure 9 According to some embodiments of this application, a schematic diagram of image retrieval is shown;
[0037] Figure 10A According to some embodiments of this application, a process for quantifying the features of an image to be retrieved is shown;
[0038] Figure 10B A schematic diagram of clustering is shown according to some embodiments of this application;
[0039] Figure 11 A schematic diagram of a Vino cell is shown according to some embodiments of this application;
[0040] Figure 12 According to some embodiments of this application, a schematic diagram of a search result is shown;
[0041] Figure 13 A schematic diagram of residual quantization is shown according to some embodiments of this application;
[0042] Figure 14 According to some embodiments of this application, schematic diagrams of clustering using different clustering methods are shown;
[0043] Figure 15 A schematic diagram of the structure of an electronic device is shown according to an embodiment of this application. Detailed Implementation
[0044] The illustrative embodiments of this application include, but are not limited to, retrieval methods, readable storage media, program products, and electronic devices.
[0045] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0046] It should be noted that the electronic devices in the embodiments of this application may also be referred to as terminals, user terminals, mobile terminals, user equipment (UE), terminal devices, mobile stations, mobile terminals (MT), etc. Terminal devices can be mobile phones, smart TVs, wearable devices, tablets, computers with wireless transceiver capabilities, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, wireless terminals in industrial control, wireless terminals in self-driving vehicles, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, etc. The following description will use mobile phones as an example. However, it is understood that the technical solutions described in this application are applicable to various electronic devices with mobile communication functions, and are not limited to mobile phones.
[0047] As mentioned earlier, a single modality can provide limited information when retrieving images. Therefore, the range of images retrieved through a single modality is relatively large, and it is not possible to accurately retrieve the image that the user wants.
[0048] The following describes a process for retrieving images.
[0049] For example, Figure 1 This diagram illustrates the process by which an electronic device retrieves images via text.
[0050] Reference Figure 1 The image retrieval interface 10 of the electronic device 100 is provided with a search box 11, in which the user can enter corresponding text to retrieve images. For example, the user can enter "trees" in the search box 11 using the virtual keyboard 20, and then click the search button 21 to retrieve the corresponding images. Then, in response to the user clicking the search button 21, the electronic device 100 can retrieve images that match the text "trees" entered by the user in the search box 11, and display the corresponding search results on the image retrieval interface 10 (for example, the search results are the images within the dashed box 12).
[0051] Continue to refer to Figure 1The detection results of electronic devices based on "trees" include images containing images of trees, images containing the words "trees," "tree," and "wood," as well as images containing descriptions of trees in other languages, such as images containing the English words "woods" or "forest." Figure 1 (Not shown in the image). Therefore, users need to continue searching for the desired image in the search results.
[0052] In some embodiments, some image retrieval models can retrieve images by inputting multimodal information such as images and text, thereby improving the accuracy of image retrieval.
[0053] For example, Figure 2 According to some embodiments of this application, a schematic diagram of image retrieval via multimodal methods is shown.
[0054] refer to Figure 2 The electronic device 100 is equipped with an image retrieval model. It can take images and text as input to the model, and then retrieve images related to the text and images based on the model. It can be understood that the electronic device 100 can store a large number of images, and users can use the image retrieval model to retrieve the desired images from this large dataset.
[0055] For example, if the input image A contains a tree, and the input text is "a person is next to a tree," the image retrieval model can directly concatenate the information provided by the text modality and the image modality to obtain the concatenated information (e.g., the concatenated information only contains information about the "tree" in image A, and information about the "person" and "tree" in the text). Then, it can retrieve the image based on this concatenated information. In other words, the subject of the images in the retrieval results output by the image retrieval model can include both the "tree" in image A and the "person" in the text. For example, in the retrieval results, image B1 is a forest containing the tree from image A, with two people playing in the forest; image B2 is an image of a person passing by a tree; image B3 is an image of a person walking on a road with trees from image A beside it; image B4 is an image of a person next to a tree, and so on.
[0056] It is understandable that while the aforementioned multimodal model can improve retrieval accuracy, the output of this model only includes images containing subjects from multiple modalities. For example, image A in the image modality includes the subject of a tree, while the text modality includes the subjects of "tree" and "person." However, the user wants to retrieve an image of "a person next to a tree," meaning there is a relative positional relationship between the "person" and the "tree." Images B1 through B4 output by the electronic device only contain the subjects of "person" and "tree," so the user still needs to search for the desired image, such as image B4, within the search results. Furthermore, if the electronic device 100 stores a large number of images, many will still meet the multimodal input criteria, causing the user to still need to spend a considerable amount of time retrieving the desired image.
[0057] In summary, image retrieval accuracy based on a single modality input is poor, while multimodal models do not consider the relationships between subjects in different modalities (e.g., relative positional relationships, which can also be referred to as subject information below). Therefore, the output of multimodal models still cannot meet the needs of users for image retrieval.
[0058] To address the aforementioned issues, this application proposes a retrieval method. After detecting a retrieval request encompassing multiple modalities, the electronic device can first determine the subject information (e.g., relative positional relationships, shape, actions, expressions, and number of target subjects) corresponding to the target subjects (e.g., people, animals, landscapes in image data, nouns in text data) in each modal of retrieval data. Then, the electronic device can determine the weight of each modal of retrieval data based on the amount of subject information provided by the target subjects in each modal; for example, the more subject information provided by each modal, the greater the weight of the corresponding retrieval data. Next, the electronic device can fuse the features of the retrieval data corresponding to each modality based on the weights of the retrieval data for each modality to obtain a first retrieval feature, and use the objects whose object features match the first retrieval feature as the retrieval results for each modality.
[0059] For example, for Figure 2In the retrieval scenario shown, the image modality retrieval data image A provides only one piece of subject information: "trees." The text modality retrieval data "a person is next to a tree" provides three pieces of subject information: "person," "tree," and the relative position of the "person" and "tree" ("person" is next to the "tree"). That is, the text modality retrieval data provides more subject information, and its weight can be set higher. In the process of fusing the text feature of "a person is next to a tree" and the image feature of image A into the first retrieval feature, because the text feature has a higher weight, the information provided by the corresponding text feature dominates, resulting in more information in the text modality retrieval data included in the fused first retrieval feature. Therefore, when performing image retrieval based on the first retrieval feature, the accuracy of image retrieval can be improved. For example, for... Figure 2 In the scenario described, using this solution, image B4 will be ranked first in the search results.
[0060] Through the above scheme, electronic devices can set higher weights for retrieval data that contains more target subjects and subject information. Therefore, the first retrieval feature can contain more target subjects and subject information, and when retrieving images based on the first retrieval feature, images that better meet the retrieval requirements can be retrieved.
[0061] In some embodiments, a search request may be a search for images and / or videos stored in an electronic device, a search for images and / or videos stored in one or more other electronic devices (e.g., a server), or a search for images and / or videos stored in an electronic device, as well as images and / or videos stored in one or more electronic devices. No limitation is made here. For ease of description, the range of images and / or videos to be searched in the search request will be referred to as the search object set.
[0062] In some embodiments, the electronic device may also store object features corresponding to images and / or videos in the search object set. Thus, during image and / or video retrieval, the electronic device does not need to perform real-time encoding of the images and / or videos in the search object set to obtain object features, thereby improving retrieval speed. Alternatively, the object features corresponding to images and / or videos in the search object set can be stored in a server or other device. The electronic device can send the first search features to the server or other device for retrieval and then send the image corresponding to the search result to the electronic device. The following uses retrieving images from a search image set as an example to illustrate the retrieval method in this application embodiment.
[0063] In some embodiments, the modality of the retrieved data may include a text modality, an image modality, an audio modality, and a video modality. For example, processing of the video modality may involve extracting keyframes from the video and converting them into an image modality. In the audio modality, processing may be performed separately, or audio may be recognized and converted into text, thereby converting the audio modality into a text modality.
[0064] In some embodiments of this application, the processing of each retrieval data may involve inputting each retrieval data modally into a corresponding model to obtain corresponding encoded data. For example, for image modal retrieval data, the retrieval data can be input into an image encoder to obtain image encoding. For text modal retrieval data, the retrieval data can be input into a text encoder to obtain text encoding. For audio modal retrieval data, the retrieval data can be input into an audio encoder to obtain audio encoding. After obtaining the encoded data of each retrieval data, the encodings of each retrieval data can be fused into a single encoding dimension to obtain a first retrieval feature. The process of fusing the encodings of multiple retrievals into a first feature may, for example, involve representing the encodings of each retrieval data in the form of a Gaussian probability density function, and then fusing the encodings corresponding to each retrieval data into a first retrieval feature based on the calculation method of the Gaussian probability density function.
[0065] In some embodiments, during the input of various search data, the target subject contained in each search data can be determined first, wherein the target subject is the subject corresponding to the image that the user needs to retrieve.
[0066] For example, after an electronic device acquires retrieval data in a text modality, it can use a text encoder to identify the various subjects within the text. This is understandable, as the text modality represents user-defined input retrieval data. Therefore, by default, users will not input text unrelated to the retrieved image; thus, all subjects provided in the text modality retrieval data can be identified as target subjects. After obtaining the data for each target subject, the text encoder also needs to determine the subject information for each target subject.
[0067] For example, if the retrieved data for the text modality is: "an image of a person playing with a cat under a tree," then the text encoder needs to determine the subject data including "person," "tree," and "cat." Furthermore, the determined subject information includes "person under a tree," "person playing with a cat," and "playing."
[0068] It is understandable that during the process of subject recognition and subject information recognition, the encoders corresponding to each modality can output the accuracy of the subject information provided by each retrieved data. Then, based on the corresponding accuracy threshold (as an example of a confidence threshold), it is determined whether to retain the subject information. In some embodiments, the accuracy threshold can be any value between 80% and 100%. It is understandable that since each encoder may be able to identify a large amount of subject information from the retrieved data, and some subject information may not be accurate enough, it is necessary to filter the subject information corresponding to each retrieved data. For example, taking an accuracy threshold of 90% as an example, [the following is an example of filtering the subject information]. Figure 2 When identifying the subject information in image B4, the accuracy of identifying the subject information "a person next to a tree" is 95%, and the probability of identifying the subject information "a person walking" is 40%. Therefore, the subject information "a person walking" can be removed.
[0069] Therefore, the subject information output by the encoder in each modality is identified in the form of accuracy. For example, in the text modality mentioned above, the accuracy of "a person under a tree" can be determined to be 100% (the text contains detailed subject information). In the image modality, the accuracy of the image encoder in recognizing subject information in the image may be biased. Therefore, it is necessary to set an accuracy threshold to filter subject information with higher accuracy.
[0070] In some embodiments, the method for determining each target subject in the retrieved data of an image modality may be, for example, to identify all subjects in the image based on a model, and then determine the target subjects to be retained in the image based on user operations. For example, after the electronic device identifies all subjects in the image, it can identify each subject individually (e.g., display the outline of each subject). The user can select the corresponding subject as the target subject. The method for selecting each target subject is, for example, clicking to select the corresponding subject or subject identifier, or deleting subjects other than the target subject based on corresponding gestures (e.g., drawing an X or a slash to indicate deletion). Alternatively, the target subject can be determined or unwanted subjects can be deleted based on voice.
[0071] The retrieval method provided in the embodiments of this application is described below.
[0072] For example, Figure 3 According to some embodiments of this application, a flowchart of an implementation method for a retrieval method is shown.
[0073] It should be noted that the electronic devices in the embodiments of this application can also be referred to as terminals, user terminals, mobile terminals, user equipment (UE), terminal devices, mobile stations, mobile terminals (MT), etc. Terminal devices can be mobile phones, smart TVs, wearable devices, tablets, computers with wireless transceiver capabilities, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, etc. The following description will use a mobile phone as an example. The executing entities of each process in the following embodiments are all electronic devices in the embodiments of this application; however, no limitation is made on the executing entity of each process when describing them.
[0074] like Figure 3 As shown, the process includes:
[0075] S301, a retrieval request including retrieval data of multiple modalities is detected, and the subject information of each target subject and / or subject of each subject in the retrieval data of each modality is determined.
[0076] For example, in some embodiments of this application, the retrieval request detected by the electronic device may be, for example, a retrieval request for an image based on retrieval data in multiple modalities. The retrieval data in multiple modalities includes, for example, image modal retrieval data, text modal retrieval data, audio modal retrieval data, and video modal retrieval data.
[0077] After detecting a search request, the electronic device can first determine the target subject in the search data of each modality. The target subject is the main subject in the image the user wants to retrieve. The target subject can be determined by the electronic device based on the user's selection or by the encoder of the corresponding modality. After determining the target subject, the subject information can be determined based on the encoder of each modality.
[0078] For example, in text modalities, a text encoder can identify individual nouns in the text to determine the target subject. It is understood that since text modalities are user-defined input retrieval data, all subjects in the user-input text modal retrieval data can be assumed to be the target subject. After identifying the target subject, the semantics of the text modal can be extracted using a text encoder to obtain the subject information provided by the corresponding retrieval data. For example, in some embodiments of this application, for text modal retrieval data, specific event or factual information can be extracted from the natural language text of the retrieval data using information extraction or event extraction models. This allows for the classification, extraction, and reconstruction of the content of the text modal retrieval data to determine the target subject and subject information provided by the text modal retrieval data. Specifically, the process of extracting the target subject and subject information from the text modal is described in detail below.
[0079] For image modalities, various subjects in an image can be identified using models that recognize image subjects, and then presented to the user for selection. Users can select a target subject or delete non-target subjects through voice, actions, gestures, clicks, etc., to determine the target subject in the image modal retrieval data. After determining the target subject provided in the image modal retrieval data, subject information for each target subject in the image can be obtained using an image encoder. Specifically, the process of extracting target subjects and subject information from image modalities is described in detail below.
[0080] For audio modalities, an audio encoder can identify the audio representing the subject. Since the audio is also user-defined input retrieval data, it can be assumed that all subjects in the input audio modal retrieval data are target subjects. After determining the target subjects provided by the audio modal retrieval data, the semantics of the audio modal retrieval data can be understood based on the audio encoder, thereby obtaining the subject information of each target subject in the audio. For example, in some embodiments of this application, when determining the target subject and subject information from the audio, the audio can be converted into text, and then the target subject or subject information can be identified through a text modal encoder. In other embodiments, features can be extracted and learned from the audio signal using models such as convolutional neural networks and recurrent neural networks, enabling the identification of different target subjects in the audio and the events corresponding to the target subjects, thereby determining the subject information. Specifically, the process of extracting target subjects and subject information from the audio modality is described in detail below.
[0081] For video modalities, keyframes can be extracted from the video modal retrieval data. These keyframes can then be processed in the same way as image modal retrieval data to obtain the target subject and subject information provided by the video modal retrieval data. For example, the video can be segmented based on the similarity of image frames. Then, corresponding keyframes can be extracted from each segment, and finally, these keyframes can be filtered to obtain keyframes for video analysis. These keyframes are then processed in the same way as image modal retrieval data to obtain the target subject and subject information provided by the video modal retrieval data. It can be understood that a video can provide multiple keyframes, and each keyframe can be considered a piece of retrieval data.
[0082] In some embodiments of this application, retrieval data from multiple different modalities can be converted into retrieval data from the same reference modality, and the subject and subject information can be extracted separately to obtain the target subject and subject information corresponding to each modality.
[0083] For example, the text modality can be used as a reference modality. Image modality retrieval data can be converted into text describing the images, and audio modality retrieval data can be converted into corresponding text, thus enabling modal recognition of subject information within the text. It can be understood that identifying the target subject and subject information within the same modality has similar recognition benchmarks, making the weighting of each retrieval data point more accurate.
[0084] In some embodiments of this application, since the retrieval data for the text modality is user-input retrieval data, it can be assumed that the target subjects provided in the retrieval data corresponding to the user-input text are relatively comprehensive. Therefore, the target subjects provided in the retrieval data corresponding to the text modality can be used as reference target subjects. When the electronic device obtains retrieval data for other modalities, it can compare the subjects identified from other modalities with the reference target subjects, thereby using the subjects identical to the reference target subjects as the target subjects in the retrieval data of other modalities, and then further identify the subject information corresponding to the target subjects of other modalities.
[0085] For example, a text dataset might provide reference subjects including "person" and "dog." An electronic device might retrieve subjects from an image including "person," "dog," "bird," and "tree." The electronic device can then use "person" and "dog" as the target subjects of the image and determine the corresponding subject information.
[0086] S302, the weight of the retrieval data for each modality is obtained based on the number of target subjects and / or the amount of information about each subject in each modality data.
[0087] For example, in some embodiments of this application, the weight of each retrieval data is related to the number of target subjects and / or the amount of subject information provided by each retrieval data.
[0088] For example, consider the amount of subject information. In a certain image retrieval process, three sets of search data are provided. The first set is in text mode, providing 3 pieces of subject information; the second set is in image mode, providing 4 pieces of subject information; and the third set is also in image mode, providing 3 pieces of subject information. In this case, the weight of the first query semantic could be 0.3, the weight of the second query semantic could be 0.4, and the weight of the third query semantic could be 0.3.
[0089] In other embodiments, the weight of each retrieval data can also be related to the amount of at least some subject information corresponding to the retrieval data and the confidence level of the subject information. For example, if there are two retrieval data, and the first retrieval data provides one subject information with a confidence level of 0.9, then 0.9 can be used as the contribution value for determining the weight of this retrieval data. The second retrieval data provides two subject information, with the first subject information having a confidence level of 0.8 and the second subject information having a confidence level of 0.9, then 0.8 + 0.9 = 1.7 can be used as the contribution value of this retrieval data. Therefore, the weight of the first retrieval data is 0.9 / (0.9 + 1.7) = 0.346, and the weight of the second retrieval data is 1.7 / (0.9 + 1.7) = 0.654.
[0090] It is understood that in other embodiments, the weight of each search data can be set based on the number of target subjects provided by each search data. Alternatively, the weight of each search data can be determined based on the sum of the number of target subjects and subject information provided by each search data. Or, after determining the number of target subjects and the number of subject information, the weights can be obtained separately based on the number of target subjects and the number of subject information. That is, a single search data can have its corresponding weight for the target subject and its corresponding weight for the subject information determined separately. Then, the weights corresponding to the target subject and the subject information are weighted and summed to obtain the weight corresponding to the search data.
[0091] For example, in some embodiments of this application, the image encoder can identify the confidence level of the subject information of each target subject in the image. The confidence level is used to represent the accuracy of the subject information, such as the accuracy of the relative position information of each target subject, the accuracy of the action of each target subject, etc. Subject information with a confidence level exceeding a confidence threshold (the confidence threshold can be set to any value between 80% and 100%) can be used as subject information for determining the weight of the image. For example, if the image encoder outputs an image corresponding to 10 pieces of subject information in the retrieval data, and 6 of them exceed the confidence threshold, then there are 6 pieces of subject information corresponding to that image.
[0092] Similarly, for text modalities, the confidence level of each subject information in the retrieved data can be output by a text encoder, thereby determining the number of subject information items in the text modal retrieval data. For audio modalities, the confidence level of each subject information in the retrieved data can be output by an audio encoder, thereby determining the number of subject information items in the audio modal retrieval data.
[0093] After determining the target subject and / or subject information of each search data, the weight data of each search data can be determined.
[0094] S303, based on the weights of the retrieval data for each modality, the features of the retrieval data corresponding to each modality are fused to obtain the first retrieval feature.
[0095] For example, in some embodiments of this application, the retrieval data for each modality can be output with corresponding feature codes by the corresponding encoder.
[0096] For example, for image modality retrieval data, the retrieval data can be input into an image encoder to obtain image encoding. Exemplarily, in some embodiments of this application, the image encoder may be, for example, a residual network (ResNet), a contrastive language-image pre-training (CLIP) model, or a CLIP-vit model that combines a CLIP model with a vision transformer (ViT) model, etc.
[0097] For text-modal retrieval data, the retrieval data can be input into a text encoder to obtain text encoding. In some embodiments of this application, the text encoder may be, for example, a global vectors forword representation (Glove) model or a gated recurrent unit (GRU) network module.
[0098] For audio modality retrieval data, the retrieval data can be input into an audio encoder to obtain audio encoding. In some embodiments of this application, the audio encoder can be, for example, a natural language processing model, a neural network model, etc. After obtaining the encoded data of each retrieval data, the encodings of each retrieval data can be fused into a single encoding dimension to obtain the first retrieval feature.
[0099] In some embodiments of this application, the codes corresponding to multiple retrieval data can be fused using a neural network. For example, a neural network (such as a multilayer perceptron, convolutional neural network, recurrent neural network, etc.) can be used to nonlinearly fuse the codes of multiple retrieval data based on the weights corresponding to the retrieval data. The neural network can automatically learn the complex relationships between features, thereby achieving more effective fusion so that the fused first retrieval feature can correspond to the code of the image to be retrieved.
[0100] In some embodiments of this application, deep learning fusion can also be used to fuse the codes corresponding to multiple retrieval data. For example, deep learning methods (such as autoencoders, variational autoencoders, generative adversarial networks, etc.) can be used to perform end-to-end learning and fusion of the codes of multiple retrieval data based on the weights corresponding to the retrieval data. These methods can capture the deep relationships between vectors and generate more representative first retrieval features, so that the fused first retrieval features can correspond to the code corresponding to the image to be retrieved.
[0101] In some embodiments of this application, the feature codes of multiple search data can be represented in the form of a Gaussian probability density function. Then, based on the calculation method of the Gaussian probability density function, the codes corresponding to each search data are fused into a first search feature according to the weight of each search data.
[0102] For example, for any retrieved data, a feature φ m and semantic features z m , can be represented as having a mean of μ m The covariance is Σ m The Gaussian distribution, where the semantic feature z m For feature φ m The processing is used to reduce the feature φ m The spatial dimensions of the data (e.g., height, width, or depth) are preserved while retaining important feature information. For example, features φ can be processed using functions such as average pooling. m To obtain semantic features z m In some embodiments of this application, only Σ can be considered. m For the case of a diagonal matrix, it can be simply written as In specific calculations, the mean value of each feature can be calculated using formulas (1) and (2). m The covariance is Σ m (or ):
[0103] μ m =LN((z m +s(fc(attn(φ m ))))) (1)
[0104]
[0105] Where LN represents layer normalization, fc represents the linear mapping layer, s(·) is the sigmoid activation function, attn represents the self-attention mechanism module, and φ m z is the feature code corresponding to the m-th retrieved data. m Let m be the semantic feature corresponding to the m-th retrieved data.
[0106] In some embodiments of this application, the following are employed: The variance is used to ensure the stability of numerical calculations. The Gaussian probability density function, besides expressing semantic information, also represents a certain degree of semantic ambiguity. For example, if the input query is "cat," this probability distribution, in addition to representing the semantic information of "cat," can also express the appearance information of different "cats." The mean μ of the fused Gaussian probability density is then determined. m and variance Then, the first search feature corresponding to the Gaussian probability density can be determined.
[0107] In some embodiments of this application, it is assumed that S = {s1, s2, ..., s} k Given k search data points, the independent distribution of all search data points can be represented by the probability distribution p(z│S) of a first search feature. i )~N(μ i ,Σ i )} 1,…,k) For example, in some embodiments of this application, the fusion process of the first retrieval feature is described using k=2 as an example. It can be understood that the two retrieval data can be of the same modality (e.g., both are image modalities) or different modalities (e.g., one is image modal retrieval data and the other is text modal retrieval data). The product of the two Gaussian probability density functions can be expressed as equation (3):
[0108] N(z;μ1,Σ1)N(z;μ2,Σ2)=N(z;μ c ,Σ c Z (3)
[0109] Where N(z; μ1, Σ1) and N(z; μ2, Σ2) are the Gaussian probability density functions of the features s1 and s2 of the two retrieved data, respectively. c ,Σ c () is the result of fusing these two Gaussian probability density functions with a mean of μ. c The covariance is Σ cThe new Gaussian probability density function, where Z is the normalization constant, is calculated using the following formula:
[0110]
[0111] Z=N(ω1μ1;ω2μ2,ω1Σ1+ω2Σ2) (6)
[0112] Where ω1 is the weight of retrieved data s1, ω2 is the weight of retrieved data s2, Σ1 and μ1 are the covariance and mean of retrieved data s1, respectively, and Σ2 and μ2 are the covariance and mean of retrieved data s2, respectively. c and μ c These are the covariance and mean corresponding to the first search feature, respectively, and Z is the normalization constant.
[0113] For the case where k>2, the synthesized Gaussian probability density function can be derived relatively easily. In this case, the normalization constant can be expressed as follows:
[0114]
[0115] Where, ω i Z represents the weight of the i-th retrieved data. i Let be the normalization constant for the i-th retrieved data.
[0116] It is understandable that after probabilistic fusion of multimodal semantic features, the multiple search data input by the user are ultimately expressed as a semantic representation of the Gaussian probability density function corresponding to the first search feature. In other words, the mean of the Gaussian probability density function corresponding to the first search feature is μ. c The variance is Σ c .
[0117] It is understood that in other embodiments, the codes corresponding to the search features can be fused based on the weights of each search data through other feature fusion methods to obtain the first search feature. The embodiments of this application do not limit the fusion method.
[0118] S304, the images whose image features in the search image set match the first search feature are used as the search results for each modality.
[0119] For example, in some embodiments of this application, after obtaining the first search feature, the electronic device can obtain an image that matches the first search feature from the search image set as the search result.
[0120] For example, the image that matches the first retrieval feature is obtained by encoding the images in the retrieval image set based on an image encoder to obtain image features, and then determining the similarity between the first retrieval feature and the image features, and selecting the image corresponding to the image feature whose similarity exceeds the similarity threshold as the retrieval result.
[0121] In some embodiments, the electronic device may store a retrieval image set, which may contain images or image features corresponding to each image. In other embodiments, the retrieval image set may be stored in a server, and the electronic device may retrieve images or image features from the server.
[0122] In some embodiments of this application, the electronic device may also store a video set. If it is necessary to retrieve a video from the video set, at least one keyframe can be extracted from a video, and then the features of the keyframe can be obtained and used as the video features of the video. If multiple keyframes are extracted from the video, the features of the multiple keyframes can be fused to obtain the keyframes of the video. Then, the similarity between the first retrieval feature and the video features of each video in the video set can be obtained, thereby selecting the video corresponding to the video feature that exceeds the similarity threshold as the retrieval result.
[0123] It is understandable that, through the above method, electronic devices can encode the retrieval data of each modality based on the encoders of each modality and determine the target subject and / or subject information provided in each retrieval data. The weight of each retrieval data can be determined based on the number of target subjects and / or subject information provided. The feature codes of each retrieval data are fused based on their weights to obtain a first retrieval feature. Since the first retrieval feature is fused based on the weights of each retrieval data, the higher the weight, the more complete the information in the corresponding retrieval data; and the more target subjects and / or subject information provided by the retrieval data, the greater the corresponding weight. Therefore, during the fusion of various retrieval data, retrieval data providing more target subjects and / or subject information can retain more information. Thus, when retrieving images based on the first retrieval feature, more attention can be paid to the target subject and its subject information, such as the relative positional relationship, actions, and expressions of the target subject, thereby improving the accuracy of image retrieval.
[0124] S305 displays the search results.
[0125] For example, in some embodiments of this application, the electronic device can display the search results after determining the search results for the corresponding search data. In some embodiments of this application, the display order of the search results is related to the similarity between the image features corresponding to the image in the search results and the first search feature. The higher the similarity, the earlier the image in the corresponding search results appears, thereby facilitating the user to quickly select the target image from the search results.
[0126] In other embodiments, if the target of the search is a video, the search result is the corresponding video.
[0127] It is understood that in some embodiments of this application, the search results may contain a large number of images with the same degree of similarity (e.g., multiple images obtained by a user through burst shooting). Therefore, the images can also be sorted according to the subject information corresponding to the target subject, with images that match the subject information more closely appearing first.
[0128] The image retrieval system of the electronic device in the embodiments of this application is described below.
[0129] For example, Figure 4 According to some embodiments of this application, an image retrieval system is shown.
[0130] Reference Figure 4 The image retrieval system includes: a subject recognition module, an encoder module, a feature fusion module, and an image retrieval module.
[0131] The subject recognition module is used to acquire various retrieval data and identify the target subject within each retrieval data. For example, the subject recognition module can identify each noun in the text modality, thereby determining the target subject in the text modality; identify each audio element representing a subject in the audio modality, thereby determining the target subject in the audio modality; and identify each subject in the image modality, which is then selected by the user through corresponding operations (e.g., gestures, actions, voice, etc.). The method for determining the target subject from the keyframes after extracting the video modality is similar to that for the image modality, and will not be elaborated here.
[0132] The encoder module includes encoders for various modalities, such as image encoders, text encoders, and audio encoders. After receiving the retrieval data and the target entities within each retrieval data item, the encoder module can encode the retrieval data and output the entity information of each target entity within the retrieval data.
[0133] For example, for image modality retrieval data (e.g., images), an image encoder can extract the feature vector corresponding to the image modality and output the subject information of each target subject in the image. For example, in some embodiments of this application, the image encoder can output a feature vector of a preset dimension corresponding to the image, as well as the subject information of each target subject in the image, such as the relative position information, actions, and expressions of each subject. Exemplarily, in some embodiments of this application, the image encoder can identify the confidence level of the subject information of each target subject in the image. The confidence level is used to represent the accuracy of the subject information, such as the accuracy of the relative position information of each target subject, the accuracy of the actions of each target subject, etc. Subject information with a confidence level exceeding a confidence threshold (the confidence threshold can be set to any value between 80% and 100%) can be used as subject information for determining the weight of the image. For example, if the image encoder outputs 10 subject information items, and 6 of them exceed the confidence threshold, then there are 6 subject information items corresponding to the image.
[0134] Similarly, for text-based retrieval data, a text encoder can extract the feature vectors corresponding to the text modality and output the subject information of each target entity in the text. For audio modality, audio information can be translated into text information and then input into a text encoder to obtain the corresponding feature vectors and subject information. Alternatively, the audio modality can be input into a corresponding audio encoder to obtain the feature vectors corresponding to the audio modality.
[0135] It is understandable that both text modality and audio modality are modal information that is user-defined input. Therefore, it can be assumed that the subjects provided in the search data of text modality and audio modality are the target subjects, so there is no need for a process of determining the target subjects.
[0136] For video modalities, keyframes can be extracted from the video. The process of processing keyframes can be referred to the process of processing information in image modalities. This application will not elaborate on the process of processing keyframes.
[0137] For example, in some embodiments of this application, the input data for encoders of different modalities is the retrieval data of the corresponding modality, and the output data is the features of the corresponding retrieval data and the confidence level of the subject information of the corresponding retrieval data. Encoders of each modality can be trained using corresponding training samples, which include input data samples of the corresponding modality and subject information samples corresponding to the input data samples. The training objective is, for example, that at least a portion of the subject information output by the encoder with a confidence level exceeding a threshold corresponds to the subject information samples.
[0138] The feature fusion module is used to determine the weight of each retrieval data based on the subject information provided by each retrieval data, and to fuse the feature vectors of each retrieval data to obtain a unified feature.
[0139] For example, in some embodiments of this application, the weight of each retrieval data is related to the amount of subject information provided by each retrieval data. For instance, in an image retrieval process, three retrieval data are provided: the first retrieval data is in text mode and provides 3 pieces of subject information; the second retrieval data is in image mode and provides 4 pieces of subject information; and the third retrieval data is also in image mode and provides 3 pieces of subject information. In this case, the weight of the first retrieval data could be 0.3, the weight of the second retrieval data could be 0.4, and the weight of the third retrieval data could be 0.3.
[0140] In some embodiments of this application, the feature vectors of each retrieved data can be represented in the form of a Gaussian probability density function, and then the feature vectors of each retrieved data can be fused based on the calculation method of the Gaussian probability density function to obtain the fused first feature.
[0141] The image retrieval module is used to retrieve images based on the first feature output by the feature fusion module. The retrieval process, for example, determines the retrieval result based on the similarity between the feature vector of the image stored in the electronic device and the first feature.
[0142] For example, in some embodiments of this application, the electronic device can encode the stored images using an image encoder to obtain corresponding image features. That is, the electronic device can pre-store the image features corresponding to the images, and when the retrieval module retrieves images, it can quickly obtain the similarity between the first feature and each image feature, and output the retrieval results in order of similarity. In some embodiments, the image retrieval module can also sort the retrieval results based on the subject information provided by each retrieval data. That is, the retrieval results are related not only to similarity but also to subject information. For example, when the similarity to the first feature is the same, the sorting order can be determined based on the subject information of each subject in the sample image.
[0143] The following describes the process of identifying each target subject and its subject information in the retrieval data of each modality through encoders of each modality.
[0144] For example, for image modality retrieval data image s1, image s1 can be input into the corresponding image recognition model to identify the various subjects in the image. Then, the user can select to keep or delete the corresponding subjects through gestures, actions, voice, clicks, etc., to determine the target subject in the target image that the user needs to retrieve.
[0145] For example, Figure 5A According to some embodiments of this application, a schematic diagram of selecting various target subjects is shown.
[0146] Reference Figure 5A An image recognition model can identify the various subjects in image s1 and assign labels to each subject. For example, label 01 is assigned to clouds in image s1, label 02 to trees, label 03 to people, and label 04 to roads. The electronic device can detect the subject corresponding to each label in image s1 clicked by the user, thus identifying that subject as the target image. For example, the electronic device detects the tree corresponding to label 02 clicked by the user as the target subject, and the person corresponding to label 03 clicked by the user as the target subject.
[0147] In other embodiments, for example Figure 5B According to some embodiments of this application, a schematic diagram of deleting various non-target entities is shown.
[0148] Reference Figure 5B In some embodiments of this application, after the electronic device identifies the various subjects in image s1 using an image recognition model, the user can draw a slash on the corresponding subject to indicate deletion. For example, if the electronic device detects that the user draws a slash on the cloud corresponding to identifier 01, it can determine to delete the cloud subject; if it detects that the user draws a slash on the road subject corresponding to identifier 04, it can delete the road subject. Therefore, the electronic device can retain the person subject identified by identifier 03 and the tree subject identified by identifier 02.
[0149] In other embodiments, the operation of determining the target subject can also be triggered by gestures, voice, actions, etc. This application does not limit the method of selecting the target subject.
[0150] In some embodiments of this application, after determining the target subject of image s1, image s1 can be input into an image encoder to obtain the feature code corresponding to image s1.
[0151] For example, in some embodiments of this application, image s1 can be input into a ResNet (which may also be called a deep residual network in other embodiments), and the ResNet model can extract features from the input image to obtain image features φ. img And the subject information of each target entity. Then, the extracted image features φ img Perform linear mapping operation f img The semantic feature vector z of the image modality can be obtained. img , i.e. z img =f img (φ imgIt can be understood that the linear mapping operation f img For example, the features of each modality are unified. That is, the output features of each modality are made to be on a unified dimension. For example, in some embodiments of this application, z img ∈R D (D=512), that is, the semantic feature vector z of the image modality img It is a 512-dimensional vector.
[0152] It is understood that in other embodiments, other models can be used when extracting the feature vectors of image modalities. For example, the feature vectors of image modalities and the relationships between various subjects in the image can be determined by the CLIP model, CLIP-vit model, etc. The embodiments of this application do not limit the model used for image processing, nor do they limit the dimension of the image feature vector mapping. For example, in other embodiments, the feature vectors of image modalities can also be mapped to 128 dimensions, 256 dimensions, 1024 dimensions, etc.
[0153] In some embodiments of this application, the subject information may include, for example, the relative positional relationships of the various target subjects, the actions of the target subjects, facial expressions, etc. Table 1 shows the subject information in image s1.
[0154] Table 1
[0155] Serial Number Subject Information Confidence (%) 1 Person on left side of tree 100 2 Person walking 95 3 Person running 60 4 Person smiling 90
[0156] For example, in some embodiments of this application, subject information with a confidence level exceeding 80% is used as subject information for determining the weight of the retrieved data. It can be understood that Table 1 provides four pieces of subject information, of which the confidence levels of the first, second, and fourth pieces of subject information exceed 80%, therefore it can be determined that image s1 provides three pieces of subject information.
[0157] The following describes the process of determining the coding features, target subject, and subject information for text modality retrieval data.
[0158] In some embodiments of this application, each word in the text data can be encoded using the Glove model algorithm (the encoding process can be, for example, by f...). glove ( ) indicates that the text data in parentheses is the input text data. Then, the features of the text modality are determined by a bidirectional GRU, denoted as f. txt In other words, for the input text s2, after f... glove Processing yields word embeddings φ txt =f glove (s2), then φ txt After inputting into the GRU module, the features of the text modality are obtained: z txt =f txt (φtxt ), where z txt ∈R D (D=512). It can be understood that in some embodiments of this application, f txt This is a text encoder, and the dimension of the feature vector of the text modality output by the text encoder needs to be the same as the dimension of the feature vector of the image modality. For example, in the embodiments of this application, the dimension of the feature vector of the text modality is also 512-dimensional. Similarly, in the embodiments of this application, when processing text data, the text modality needs to output the relationship between each target subject in the text. It can be understood that since the text data is input by the user, it can be regarded that all the subjects in the text input by the user are target subjects, that is, the subjects in the text are no longer filtered.
[0159] For example, in the above embodiments of this application, the model for encoding text is only one example. In other embodiments, other models can be used to extract the features of the text and the relationships between the various subjects in the text. The embodiments of this application do not limit the model for processing text data.
[0160] In some embodiments of this application, information extraction or event extraction can be used to extract specific event or factual information from natural language text, thereby classifying, extracting, and reconstructing the text content to determine the target subject and subject information provided by the text modality retrieval data. This information typically includes entities, relationships, and events. For example, extracting time, location, and key figures from news articles, or extracting product names, development times, and performance indicators from technical documents. In some embodiments of this application, information extraction can be used to determine the subject information in the text modality.
[0161] For example, Figure 6 According to some embodiments of this application, a schematic diagram of text and labels is shown.
[0162] like Figure 6 As shown, the retrieval data of the text modality can be sorted according to... Figure 6 The table shown is labeled with different tags, and then the individual entities and the relationships between them are identified.
[0163] For example, relation extraction (also known as triple extraction) can be used to extract relationships between target entities. After obtaining the retrieval data of the text modality, it is possible to... Figure 6 The tags in the algorithm are used to label the retrieved data, thereby identifying the target subjects in the text-modal retrieved data, and then extracting the semantic relationships between two or more target subjects. Semantic relationships are usually used to connect two target subjects and, together with the target subjects, express the main meaning of the text-modal retrieved data.
[0164] In some embodiments of this application, the subject predication object (SPO) can be used to represent the data. For example, if the text in the search data is "apples are red", then the extracted triple can be (apple color is red), and the subject information of the search data can be determined based on the triple. Exemplarily, in some embodiments of this application, the subject information corresponding to the text modality of the search data can be determined using Table 2.
[0165] Table 2
[0166]
[0167]
[0168]
[0169] For example, referring to Table 2, subject relationships can be categorized into positional relationships, relationships between objects, biological actions, human behavior, traffic scenes, motion scenes, and background relationships. After determining the semantics of the text, the subject information of each target subject can be determined using Table 2.
[0170] It is understandable that since the retrieval data for the text modality is user-defined input, it can be assumed that all subjects contained in the text modality are the target subjects, and therefore, there is no need to filter the target subjects.
[0171] The following describes the process of determining the coding features, target subject, and subject information from the retrieved data of audio modalities.
[0172] For the input audio s3, existing speech processing techniques can be used to translate the speech information into text; then, text modality processing methods can be used for encoding to obtain the semantic representation of the speech modality. Alternatively, a closed-loop attention-based pretraining (CLAP) backbone network f can be used. clap Feature extraction is performed to obtain audio features φ audio Then, a linear mapping operation f is performed on the extracted audio features. audio This yields the feature vector of the audio modality. For example: z audio =f audio (φ audio Among them, z audio ∈R D Among them, f audioFor audio encoders, the same method is used to map the feature vectors of the audio modality to the same dimensions as the feature vectors of the image and text modalities. For example, the feature vectors of the audio modality are also 512-dimensional.
[0173] In some embodiments of this application, after the audio encoder obtains the retrieval data of the audio modality, it can extract the speech frame (audio with sound) from the audio s3, thereby removing invalid data (audio segments without sound or interfering audio) in the audio s3, and then identify the semantics in the audio s3.
[0174] For example, Figure 7 According to some embodiments of this application, a schematic diagram of a target subject in audio extraction is shown.
[0175] like Figure 7 As shown, after acquiring audio s3, the audio encoder can extract the speech segments in audio s3, and then, based on machine learning or deep learning models, identify the semantics corresponding to the speech segments, thereby determining the target subject and subject information in audio s3.
[0176] The following describes the process of determining coding features, target subjects, and subject information from video modality retrieval data.
[0177] For the input video s4, keyframes can be obtained through appropriate processing. Video contains more information than images, but a video contains too much redundant information, so it is necessary to extract the keyframes.
[0178] For example, Figure 8 According to some embodiments of this application, a schematic diagram of extracting keyframes from a video is shown.
[0179] Reference Figure 8 When extracting keyframes, the video can be divided into multiple segments based on similarity. Then, representative keyframes can be extracted from each segment. The numerous keyframes can then be evaluated (e.g., image brightness, sharpness, contrast, and saturation) to select the image content with the best overall quality.
[0180] In some embodiments of this application, keyframes in each video segment can be obtained using methods based on image content, motion analysis, trajectory curve point density features, and clustering. For example, in some embodiments of this application, the K-means algorithm can be used to extract keyframes for each video. In other embodiments, video editing tools such as audio / video codec tools (Fast Forward Moving Picture Experts Group, FFmpeg) can also be used for keyframe extraction.
[0181] After obtaining the keyframes of each video segment, a basic image quality assessment model can be used. This involves taking the image to be evaluated as input and outputting a corresponding score, such as brightness, sharpness, contrast, and saturation. Alternatively, a more in-depth aesthetic evaluation module can be used to perform finer-grained evaluations of the keyframe images, such as overall image score, lighting score, content score, background score, foreground score, and composition score. The extracted keyframes are then filtered to obtain a multi-frame image for querying. After obtaining the multiple keyframes extracted from the video, they can be processed using the aforementioned image modality retrieval data processing methods to obtain the target subject and subject information corresponding to each keyframe.
[0182] Below is a schematic diagram illustrating image retrieval based on retrieval data from multiple modalities.
[0183] For example, Figure 9 According to some embodiments of this application, a schematic diagram of image retrieval is shown.
[0184] For example, such as Figure 9 As shown, the retrieval data for the image includes image s1 in the image modality, text s2 in the text modality, audio s3 in the audio modality, and video s4 in the video modality.
[0185] For example, video s4 can be first input into the image extraction module to extract keyframes, and then the keyframes from video s4 and image s1 can be input into the subject recognition module. The subject recognition module can identify the various subjects in the keyframes and image s1, and the user can select the target subject from the keyframes and image s1 through corresponding click operations.
[0186] Then, image s1 and keyframes can be input into the image encoder to obtain the corresponding image features and image weights based on the image encoder.
[0187] Similarly, the text s2 of the text modality can be input into the text encoder to obtain text features and text weights. The audio s3 of the audio modality can be input into the audio encoder to obtain audio features and audio weights.
[0188] It can be understood that image weights, text weights, and audio weights are determined by the encoder module based on the number of target subjects and / or subject information corresponding to each retrieval data (e.g., keyframes of image s1, text s2, audio s3, and video s4) output by the encoder of each modality. After determining the number of subjects and subject information for each retrieval data, the encoder module can input the number of subjects and subject information into the image retrieval module to sort the target images output by the image retrieval module.
[0189] After obtaining the feature codes corresponding to each retrieval data, the feature fusion module can fuse the feature codes of each retrieval data based on the weights of each retrieval data. It can be understood that a keyframe provided in video s4 represents one retrieval data.
[0190] After the feature fusion module merges the feature codes of each retrieval data into a first retrieval feature, the first retrieval feature can be input into the image retrieval module so that the image can be retrieved based on the first retrieval feature.
[0191] In some embodiments of this application, the electronic device stores a retrieval image set. During the image retrieval process of the image module, the electronic device can obtain images from the retrieval image set and encode the images based on an image encoder to obtain the image features to be retrieved. The image retrieval module can determine the image features that match the first retrieval feature from the image features to be retrieved and output the target image corresponding to the image feature.
[0192] For example, in some embodiments of this application, the target image can be determined based on the similarity between the image features to be retrieved and the first retrieval features.
[0193] In some cases, there may be multiple images in the target image that have the same similarity to the first retrieval feature. In this case, the target images with the same similarity can be sorted based on the subject information of each target subject.
[0194] In other scenarios, electronic devices can also pre-store the searchable features corresponding to images in the search image set. That is, during the image retrieval module's search based on the first search feature, the electronic device no longer needs to encode the images in the search image set. In other embodiments, the search image set can also be stored on a server, and the electronic device can retrieve the corresponding images or the corresponding searchable image features from the server.
[0195] The above-described image retrieval process can adjust the weight of the feature codes corresponding to each retrieval data based on the subject information provided by each retrieval data, so that the retrieval data with more subject information can provide more information in the first retrieval feature, thereby improving the accuracy of electronic devices in retrieving images based on the first retrieval feature.
[0196] In some embodiments of this application, to accelerate the retrieval process, the electronic device may pre-store the image features to be retrieved. Then, the image features to be retrieved are quantized to improve retrieval speed.
[0197] For example, Figure 10A According to some embodiments of this application, a process for quantifying the features of an image to be retrieved is shown.
[0198] For a given feature vector, product quantization (PQ) can be used to compress the feature vector D, thereby reducing storage space and data search time.
[0199] like Figure 10A As shown, data X consists of M D-dimensional feature vectors. For example, in some embodiments of this application, the image set to be retrieved may store data X, where M represents the number of images in the image set and D represents the dimension of the images in the image set after being encoded by an image encoder.
[0200] PQ quantization can divide a high-dimensional vector into multiple smaller sub-vectors. For example, in some embodiments of this application, a D-dimensional vector can be divided into four d-dimensional sub-vectors. Thus, the data X composed of M D-dimensional vectors can be divided into four sub-data, such as sub-data X1, sub-data X2, sub-data X3, and sub-data X4.
[0201] Then, each sub-data is clustered using a clustering algorithm (such as k-means) to generate a corresponding codebook C. Each codeword in codebook C is a low-dimensional vector representing a cluster center.
[0202] Then, for each of the M sub-vectors in the sub-data, find the closest codeword in the corresponding codebook C, and use the cluster center of that codeword to represent the sub-vector.
[0203] For example, in some embodiments of this application, the sub-data can be divided into N classes, and N sub-vectors can be randomly selected as initial cluster centers. Then, M sub-vectors are assigned to the classes closest to each cluster center. The cluster centers in each class are then updated. For example, the k-means algorithm can be used to take the mean of multiple sub-vectors in each class as the cluster center for that class. The M sub-vectors are then reassigned to the classes closest to each cluster center, and the cluster centers are updated again. Each update of the cluster centers can be called an iteration. Through multiple iterations, the position of the cluster centers is optimized. In each iteration, each sub-vector is assigned to the class containing the nearest cluster center. Afterward, the cluster center of each class is updated to the mean of all sub-vectors within that class.
[0204] The iterative process will continue until the location of the cluster centers no longer changes significantly, or until the preset number of iterations is reached.
[0205] It is understandable that the above clustering process could be, for example, dividing M sub-vectors in a subset of data into N classes, with each class having a cluster center, and then determining which class each M sub-vector belongs to based on the distance (or similarity) between the M sub-vectors and the N cluster centers.
[0206] For example, Figure 10B A schematic diagram of clustering is shown according to some embodiments of this application.
[0207] like Figure 10B As shown, subvector X11 is a subvector in data X1, and subvector X11 can be classified using the clustering module.
[0208] For example, in some embodiments of this application, N is 128, meaning that the M sub-vectors in a sub-data set are divided into 128 classes. Sub-vector X11 has the smallest distance to the cluster center of the n1th class, so it can be assigned to the n1th class, and thus represented by the cluster center of the n1th class. Similarly, for a sub-vector X21 in sub-data set X2, the clustering module can determine that it is assigned to the n2th class, sub-vector X31 in sub-data set X3 is assigned to the n3th class, and sub-vector X41 in sub-data set X4 is assigned to the n4th class. Here, n1, n2, n3, and n4 are all less than or equal to N. When dividing each sub-vector into N classes, the cluster center (or centroid) of each class can be determined, and then all sub-vectors in that class can be represented by the cluster center of each class.
[0209] For example, Figure 11 A schematic diagram of a Vinio cell is shown according to some embodiments of this application.
[0210] like Figure 11 As shown, cluster centers of a class can be used to represent subvectors belonging to that class, thus preserving important information about that class. For example, clustering operations can yield Voronoi cells, such as... Figure 11 Different regions in the data can be represented as a category, the centroid represents the cluster center, and the sample features represent the sub-vectors belonging to that category. It can be understood that similar sub-vectors are assigned to the same region (or cell), meaning similar sub-vectors belong to the same category, and then the sub-vectors of that category are represented by the centroid (cluster center) of that region.
[0211] In some embodiments of this application, sub-vectors X11, X21, X31, and X41 can be, for example, four sub-vectors divided from a D-dimensional vector. This D-dimensional vector can then be represented as (n1, n2, n3, n4).
[0212] When retrieving data from data X using retrieval features, the retrieval features can also be divided into 4 sub-retrieval data segments. Then, each sub-retrieval data segment is classified to determine its category. All sample features within that category are similar to those of the sub-clue data.
[0213] The following describes the process of retrieving data.
[0214] Once the codebook C is determined, data can be retrieved quickly using an inverted file. The inverted file accelerates searches; for each feature vector, it stores a list of data containing that feature vector, such as codebook C. This allows for rapid location of data containing similar features during a query.
[0215] For example, for a query vector K, similar data can be retrieved using asymmetric distance. For instance, following the process of generating codebook C from data X, it can be divided into identical sub-segments, such as sub-vectors K1, K2, K3, and K4. Then, the distances from sub-vectors K1 to K4 to the N cluster centers corresponding to codebook C are calculated, resulting in 4×N distances. For ease of understanding, this 4×N can be called a distance table.
[0216] When determining the distance LG from a sample G in data X to the query vector K, the cluster position of sample G in the codebook C can be determined. For example, if the position of sample G in codebook C is (98, 56, 124, 13), then the distances L1 from sub-vector K1 to the 98th cluster, L2 from sub-vector K2 to the 56th cluster, L3 from sub-vector K3 to the 124th cluster, and L4 from sub-vector K4 to the 13th cluster can be determined from the distance table. Then, the distances from each sub-vector to sample G are summed to determine the distance LG between sample G and the query vector K, i.e., LG = L1 + L2 + L3 + L4. After determining the distances between all samples and the query vector K, they can be sorted by distance to identify the sample in data X most similar to the query vector K. This improves retrieval speed.
[0217] It is understood that in the embodiments of this application, the query vector K may be the vector corresponding to the first retrieval feature.
[0218] For example, in other embodiments, data can also be retrieved through other retrieval methods, such as symmetrical distance, etc. The embodiments of this application do not limit the process of retrieving data.
[0219] In some embodiments of this application, since the electronic device's image library may contain multiple highly similar images—for example, a user might take a series of photos when capturing a memorable moment—a large number of highly similar images will be captured. During the retrieval process, the electronic device may output a large number of highly similar images.
[0220] For example, Figure 12 A schematic diagram of a search result is shown according to some embodiments of this application.
[0221] like Figure 12 As shown, the electronic device detected that the user's input search data included the text "person jumping over a stream" and image a. The output search results were images a1 to a6, all of which depicted a person jumping over a stream. The user obtained six highly similar images using burst mode. Therefore, after the electronic device output the search results, the user still needed to select the desired image from the multiple highly similar images.
[0222] Therefore, in some embodiments of this application, when retrieving data based on product quantization, the image data stored in the image retrieval center can be further processed to improve retrieval accuracy.
[0223] For example, Figure 13 A schematic diagram of residual quantization is shown according to some embodiments of this application.
[0224] It should be noted that, Figure 13 In this embodiment, the feature processing method for images in the retrieved image set can be referred to Figure 10A The product quantization in the embodiments, and Figure 10A The difference in the implementation is that after obtaining the codebook C, it is also necessary to obtain the residual R between the data X and the corresponding cluster center in the codebook C.
[0225] like Figure 13 As shown, after obtaining the codebook C based on data X, the residual R between data X and codebook C can be obtained. For example, taking sub-residual R1 as an example, R1 is the residual between the M sub-vectors in sub-data X1 and the cluster centers of the corresponding categories in codebook C. For sub-vector X11 in sub-data X1, its residual sub-vector R11 = X11 - N1, where N1 is the cluster center of sub-vector X11 and the n1th class. That is, the number of residual sub-vectors in sub-residual R1 is the same as the number of sub-vectors in sub-data X1, both being M. Correspondingly, the data of sub-residuals R2 to R4 can be obtained.
[0226] After obtaining the sub-residuals R1 to R4, they can be further clustered using corresponding clustering algorithms. For example, in some embodiments of this application, the k-median clustering algorithm can be used for clustering. After determining the cluster centers of each category using the k-mediods algorithm, the nearest actual sample point to this cluster center can be selected as the centroid. That is, k-mediods uses the actual sample points in the residual dataset as its centroids.
[0227] For example, Figure 14 According to some embodiments of this application, schematic diagrams of clustering using different clustering methods are shown.
[0228] like Figure 14 As shown, when updating cluster centers using the k-means method, the mean of a category is used as the cluster center. However, when updating cluster centers using the k-mediods method, the nearest sub-vector to the specified sub-vector is used as the cluster center. In some embodiments, the mean of a category can be determined first, and then the sub-vector closest to the mean can be used as the cluster center. k-mediods is suitable for scenarios that are sensitive to noise and outliers. Due to its strong robustness to outliers, the k-mediods algorithm is more advantageous when dealing with datasets containing noise and outliers.
[0229] The training process of the feature fusion module in the embodiments of this application is described below.
[0230] For example, when performing feature fusion and feature retrieval, the feature encoding extracted from the retrieval data of each modality needs to undergo dimensionality transformation through feature mapping. A certain amount of training data is required to fine-tune each encoder to ensure that the fused first retrieval feature matches the retrieved result. Therefore, the feature fusion process requires training with corresponding training data.
[0231] The training data consists of various search conditions and returned result images, such as triples (s1, s2, t1), where s1 and s2 are a set of image-image data pairs, text-text data pairs, or image-text data pairs with similar semantics, and t1 is the corresponding query result. The training data at this stage is mainly generated manually through annotation. In other embodiments, combinations of more search data and result images can also be used as training data, such as audio modality search data, image modality search data, text modality search data, and result images. The embodiments of this application do not limit the number of modalities in the training data or the number of search data. The goal of model training is to make the Gaussian probability density distribution of the fused feature vector as close as possible to the Gaussian probability distribution of the target image, while simultaneously making the Gaussian probability density distribution of the fused feature vector as far away as possible from the Gaussian probability distribution of the non-target image.
[0232] To measure the similarity between two probability distributions, J data points can be sampled from the two distributions p(z│x) and p(z│y). Generate J 2 The similarity between two distributions can be calculated using data pairs, for example, by referring to formula (8):
[0233]
[0234] Where p(z│x) and p(z│y) are the Gaussian probability density distributions of the feature codes of the two retrieved data. k(·,·) is a standard metric function, such as cosine similarity. J is the number of sampling points. and These are the sampling points for the Gaussian probability density of the two retrieved data.
[0235] In the specific calculation, since the feature fusion results in a multivariate Gaussian distribution p(z│S), and the distribution of the resulting image is p(z│t), for ease of calculation, we can take its logarithm, i.e., transform it as follows:
[0236]
[0237] Where p(z│t) is the Gaussian probability density distribution of the resulting image. p(z│S) is the Gaussian probability density distribution after fusion. Z is the normalization constant, and the objective function of the network learning is equation (10):
[0238]
[0239] Where B is the batch data during model training, i.e., the number of triples (s1, s2, t1). To ensure the stability of training, a regularization term is usually added, such as in equation (11):
[0240]
[0241] Where |S| is the number of all input query conditions. It is the variance of the j-th input and the i-th query pair. In summary, the final objective function is equation (12):
[0242]
[0243] The feature fusion module can be trained using equation (12) to ensure that the first retrieval feature after fusion can match the feature corresponding to the retrieval image, thereby improving retrieval accuracy.
[0244] The following section uses a mobile phone as an example to provide a detailed description of the electronic devices involved in some embodiments of the present invention.
[0245] Figure 15 A schematic diagram of the structure of an electronic device is shown according to an embodiment of this application.
[0246] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0247] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0248] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.
[0249] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.
[0250] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0251] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0252] The charging management module 140 is used to receive charging input from the charger.
[0253] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, display 194, camera 193, and wireless communication module 160, etc.
[0254] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0255] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization.
[0256] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1.
[0257] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through audio devices (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.
[0258] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0259] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a glass cover 10 and a display panel 20. The display panel 20 can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini-LED, a Micro-LED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0260] Camera 193 is used to capture still images or videos.
[0261] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.
[0262] Internal memory 121 can be used to store computer executable program code, which includes instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of electronic device 100 by running instructions stored in internal memory 121 and / or instructions stored in memory located in the processor.
[0263] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0264] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.
[0265] This application also provides a program product that, when executed on an electronic device, enables the electronic device to implement the methods provided in the foregoing embodiments.
[0266] This application also provides a readable storage medium storing one or more programs, which, when executed by an electronic device, enable the electronic device to implement the methods provided in the foregoing embodiments.
[0267] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. Embodiments of this application can be implemented as computer programs or program code executable on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0268] Program code can be applied to input instructions to execute the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a digital signal processor, a microcontroller, an application-specific integrated circuit, or a microprocessor.
[0269] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this application are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0270] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried on or stored thereon on one or more transient or non-transitory machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or through other computer-readable media. Therefore, machine-readable media can include any mechanism for storing or transmitting information in a machine-readable (e.g., computer-readable) form, including but not limited to floppy disks, optical disks, CD-ROMs, compact disc-read-only memory (CD-ROMs), magneto-optical disks, read-only memory (ROM), random-access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic cards or optical cards, flash memory, or tangible machine-readable storage for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in the form of electrical, optical, acoustic, or other forms of propagation signals. Therefore, machine-readable media includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine-readable (e.g., computer-readable) form.
[0271] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Furthermore, the inclusion of structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.
[0272] It should be noted that all units / modules mentioned in the device embodiments of this application are logical units / modules. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important factor; the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. Furthermore, to highlight the innovative aspects of this application, the above-described device embodiments of this application have not introduced units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that the above-described device embodiments do not contain other units / modules.
[0273] It should be noted that in the examples and description of this patent, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0274] Although this application has been illustrated and described with reference to certain preferred embodiments thereof, those skilled in the art should understand that various changes in form and detail may be made thereto without departing from the spirit and scope of this application.
Claims
1. A retrieval method applied to electronic devices, characterized in that, include: The system detects retrieval requests that include retrieval data in multiple modalities and determines at least some subject information corresponding to the target subject in each retrieval data. Based on the weights of each search data, feature fusion is performed on the features corresponding to each search data to obtain a first search feature, wherein the weights of each search data are related to the features of the at least part of the subject information corresponding to each search data, and the features of the at least part of the subject information include the quantity of the at least part of the subject information; In the search object set, objects whose object features match the first search feature are used as the search results for each of the search data. The greater the weight of the corresponding search data, the greater the relevance between the search result and the search data.
2. The method according to claim 1, characterized in that, The mode includes at least one of the following modes: Image modality, text modality, audio modality, video modality.
3. The method according to claim 1, characterized in that, The subject information includes at least one of the following: The relative positional relationships of the target subjects, the inter-object relationships of the target subjects, the biological actions of the target subjects, the behavior of the target subjects, the traffic scene in which the target subjects are located, the movement scene of the target subjects, the relationship between the target subjects and the background, and the number of target subjects.
4. The method according to claim 1, characterized in that, The determination of at least some subject information corresponding to the target subject in each retrieved data includes: The confidence level of the subject information of the target subject in each of the retrieved data is obtained, and the confidence level is used to indicate the accuracy of the corresponding subject information; The subject information with a confidence level greater than a confidence threshold among the subject information of each of the retrieved data shall be used as the at least part of the subject information of the corresponding retrieved data.
5. The method according to claim 1, characterized in that, The step of determining at least some subject information corresponding to the target subject in each of the retrieved data further includes: The retrieved data is converted into the transformed retrieved data corresponding to the reference modality; Determine the transformation subject information corresponding to the target subject in each transformation retrieval data; The transformed subject information is used as at least part of the subject information of the corresponding retrieval data.
6. The method according to claim 1, characterized in that, The weight of each retrieval data is the proportion of the number of at least part of the subject information corresponding to each retrieval data to the number of at least part of the subject information of all retrieval data.
7. The method according to claim 1, characterized in that, The search object set includes an image set and / or a video set.
8. The method according to claim 1, characterized in that, The determination of at least some subject information corresponding to the target subject in each retrieved data includes: Corresponding to the retrieved data being an image modality, at least one subject in the image of the retrieved data for the image modality is identified; Based on user actions, at least a portion of the at least one subject is selected as the target subject corresponding to the search data; Based on the target subject in the image, at least a portion of the subject information of the retrieved data for the image modality is determined.
9. An electronic device, characterized in that, Includes: memory, used to store instructions; At least one processor is configured to execute the instructions to cause the electronic device to implement the method of any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the method of any one of claims 1 to 8.
11. A computer program product, characterized in that, When the computer program product is run on the device, it causes the device to perform the method of any one of claims 1 to 8.