Searching method and device, model training method and device and storage medium

CN120129901APending Publication Date: 2025-06-10HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280101313.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-10-28
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing on-device search technology cannot perform concrete searches, resulting in poor search results and poor user experience. Furthermore, due to resource constraints and security and privacy factors, cloud-side search capabilities cannot be used.

Method used

By receiving the user's search request, the characteristics of the private entity are determined, and the search results are obtained based on the characteristics to achieve the correspondence between the private entity and the actual image. The model training method is used to generate fusion features to support concrete search, without the need for closed tag sets, and is flexible Strong flexibility and adaptable to diversified search requests.

Benefits of technology

It improves the search effect and efficiency, improves the user experience, realizes concrete search, is suitable for different terminal devices, does not need to change models, and is easy to maintain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120129901A_ABST
    Figure CN120129901A_ABST
Patent Text Reader

Abstract

The invention relates to a search method, a model training method and device and a storage medium. The search method can comprise the steps that a first search request of a user is received, and the first search request comprises a private entity; in the features corresponding to the at least one private entity, determining the features corresponding to the private entity in the first search request; wherein the feature corresponding to the at least one private entity indicates data corresponding to the at least one private entity, and the data corresponding to the at least one private entity and the first search request correspond to different modals; obtaining a search result according to the first search request and the features corresponding to the private entities in the first search request; the search result and the first search request correspond to different modes; displaying a search result; through the application, the problem that the terminal equipment cannot perform concrete search in the search process can be solved, and the search effect and the search efficiency are improved, so that the search experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

A search method, model training method, device and storage medium Technical Field

[0001] The present application relates to the field of computer search technology, and in particular to a search method, a model training method, a device and a storage medium. Background Art

[0002] Currently, mainstream search engines in the industry are all cloud-oriented. Search capabilities based on cloud-based services have made significant progress. Indexing technologies include both traditional word-based inverted indexes and the popular approximate nearest neighbor (ANN) index based on model-based inference vectors. However, on mobile phones, tablets, and other terminal devices, various resource limitations (such as memory, storage, and power consumption) prevent cloud-based search practices from being migrated to the client side. These limitations require significant tailoring, and the final results are difficult to guarantee. Furthermore, due to security and privacy concerns, cloud-based capabilities cannot be used to search client-side data.

[0003] However, existing end-side search technologies cannot perform concrete searches, that is, they cannot match entities in the query text with specific images in reality, resulting in poor search results and a poor user search experience.

[0004] Summary of the Invention

[0005] In view of this, a search method, a model training method, a device, an electronic device and a storage medium are proposed.

[0006] In a first aspect, an embodiment of the present application provides a search method, comprising: receiving a first search request from a user, the first search request including a private entity; wherein the private entity represents an entity having an associated relationship with the user; determining, among the features corresponding to at least one private entity, the features corresponding to the private entity in the first search request; wherein the features corresponding to the at least one private entity indicate data corresponding to the at least one private entity, and the data corresponding to the at least one private entity corresponds to a different modality from the first search request; obtaining search results based on the first search request and the features corresponding to the private entity in the first search request; the search results correspond to a different modality from the first search request; and displaying the search results.

[0007] Based on the above technical solution, among the features corresponding to at least one private entity, the features corresponding to the private entity in the user's search request are determined, and the search results are obtained based on the features corresponding to the private entity and the first search request. Since the features corresponding to at least one private entity indicate the data corresponding to at least one private entity, the data corresponding to at least one private entity can correspond to the same modality as the search results, so that the private entity in the search request can be matched with the specific image in reality (i.e., the specific image of the private entity in the search results), thereby realizing a figurative search; solving the problem that the terminal device cannot perform a figurative search during the search process, improving the search effect and search efficiency, and enhancing the user's search experience. As an example, the model can be used to infer the private entity in the first search request and the first search request to obtain public-private fusion features, and then the search results can be determined in the database through the public-private fusion features. In addition, there is no need for a closed tag set, which is highly flexible; it can meet the user's diversified search requests and has higher scalability; for different terminal devices, there is no need to replace the model, and maintenance is easier.

[0008] According to the first aspect, in a first possible implementation of the first aspect, obtaining search results based on the first search request and features corresponding to private entities in the first search request includes: processing the first search request and features corresponding to private entities in the first search request through a model to generate a fusion feature; the fusion feature indicates the first search request; and using data corresponding to features in a database that match the fusion feature as the search results.

[0009] Based on the above technical solution, the model is used to infer the first search request and generate fusion features, which can simultaneously represent public information and concrete private information; as an example, the first search request can be processed by the model to generate public features, and then the generated public features and private features (i.e., the features corresponding to the private entities in the first search request) are fused to generate fusion features; thereby, for different users, the search results that the users want can be searched, and concrete search based on the fusion of private features and public features can be realized.

[0010] According to the first aspect or the first possible implementation of the first aspect, in the second possible implementation of the first aspect, the method further includes: obtaining data corresponding to a first private entity; the first private entity is any one of the at least one private entity; processing the data corresponding to the first private entity to generate features corresponding to the first private entity.

[0011] Based on the above technical solution, data corresponding to the first private entity is obtained. As an example, the data corresponding to the first private entity can be obtained by explicit collection or implicit collection, so as to collect concrete private information within the scope permitted by privacy security; and the data corresponding to the first private entity is processed to generate features corresponding to the first private entity, thereby obtaining features corresponding to the first private entity; illustratively, the features corresponding to the first private entity can be stored in the terminal device; thereby, in the process of subsequent multimodal search performed by the terminal device, support is provided for quickly determining the features corresponding to the target private entity.

[0012] According to the second possible implementation manner of the first aspect, in a third possible implementation manner of the first aspect, obtaining data corresponding to the first private entity includes: issuing a first prompt message to the user, the first prompt message being used to prompt the user to select data containing the first private entity; and obtaining the data corresponding to the first private entity in response to the user's selection operation.

[0013] Based on the above technical solution, the data corresponding to the first private entity is obtained in an explicit collection method that can be perceived by the user. Since the private information is subjectively selected by the user, the private information collected in this way has the characteristic of high confidence; thereby achieving the collection of highly confident and concrete private information through display guidance within the scope of user privacy security.

[0014] According to the second possible implementation manner of the first aspect, in a fourth possible implementation manner of the first aspect, obtaining data corresponding to the first private entity includes: obtaining initial data containing the first private entity; displaying at least one search result in response to a second search request from a user; the second search request contains the first private entity, and the at least one search result corresponds to a different modality from the second search request; the at least one search result includes the initial data; and based on the user's operation of selecting the initial data in the at least one search result, using the initial data as the data corresponding to the first private entity.

[0015] Based on the above technical solution, low-confidence concrete private information (i.e., initial data) is automatically obtained, and further in response to a second search request actively triggered by the user, at least one search result is displayed, and based on the user's operation of selecting the above initial data in at least one search result, the initial data is used as the data corresponding to the first private entity, thereby improving the confidence of the concrete private information, thereby obtaining the data corresponding to the first private entity in an implicit collection method that the user is unaware of.

[0016] According to the second, third or fourth possible implementation manner of the first aspect, in the fifth possible implementation manner of the first aspect, the method further includes: obtaining data containing the first private entity and the second private entity; wherein the second private entity is any private entity other than the first private entity in the at least one private entity; and obtaining features corresponding to the second private entity based on the data containing the first private entity and the second private entity, and features corresponding to the first private entity.

[0017] Based on the above technical solution, the features corresponding to the second private entity are obtained according to the acquired data containing multiple private entities and the generated private knowledge (i.e., the features corresponding to the first private entity); as an example, common sense and models can be used to obtain the features corresponding to the second private entity, and the private knowledge set can be automatically expanded. Therefore, within the scope permitted by privacy security, with as little collection as possible and a small amount of generated private knowledge, the concrete private knowledge can be improved, thereby realizing the effective use of data in situations such as when there is little user feedback data, and providing support for realizing concrete search in the process of terminal devices performing multimodal search.

[0018] According to the foregoing various possible implementations of the first aspect, in a sixth possible implementation of the first aspect, the private entity includes: a title, a name, or a nickname associated with the user.

[0019] In a second aspect, an embodiment of the present application provides a search method, comprising: displaying a first display interface; the first display interface includes a search portal; in response to a user inputting a keyword in the search portal, displaying a second display interface, the second display interface including a first search result corresponding to the keyword; wherein the keyword includes a private entity; the first search result is obtained based on the features corresponding to the keyword and the private entity; the first search result and the keyword correspond to different modalities.

[0020] Based on the above technical solution, in response to the user entering a keyword in the search entrance of the first display interface, the second display interface is displayed; since the first search result is obtained based on the keyword and the characteristics corresponding to the private entity in the search request, the private entity in the keyword can be matched with the specific image in reality (that is, the specific image of the private entity in the search result), thereby realizing a concrete search; this solves the problem that the terminal device cannot perform a concrete search during the search process, improves the search effect and search efficiency, and enhances the user's search experience.

[0021] According to the second aspect, in a first possible implementation of the second aspect, a third display interface is displayed, and the third display interface includes first prompt information and at least one data; the first prompt information is used to prompt the user to select data containing a first private entity; in response to the user's selection operation in the at least one data, a fourth display interface is displayed, and the fourth display interface includes first data and a first identifier, and the first identifier is used to indicate that the first data is the data selected by the user.

[0022] Based on the above technical solution, the data corresponding to the first private entity is obtained in an explicit collection method that can be perceived by the user. Since the private information is subjectively selected by the user, the private information collected in this way has the characteristic of high confidence; thereby achieving the collection of highly confident and concrete private information through display guidance within the scope of user privacy security.

[0023] According to the second aspect or the first possible implementation of the second aspect, in the second possible implementation of the second aspect, the second display interface also includes: second prompt information, the second prompt information is used to prompt the user to confirm whether the first search result is the search result expected by the user; the method also includes: in response to the user's confirmation operation, displaying a fifth display interface, the fifth display interface including the first search result and a second identifier, the second identifier is used to indicate that the first search result is the search result confirmed by the user.

[0024] Based on the above technical solution, data corresponding to private entities is obtained in an explicit collection method that can be perceived by the user. Since the private information is confirmed by the user, the private information collected in this way has the characteristic of high confidence. This enables the collection of highly confident, concrete private information through display guidance within the scope of user privacy security.

[0025] In a third aspect, an embodiment of the present application provides a model training method, comprising: obtaining a multimodal sample set and a first model; wherein the multimodal sample set comprises: a multimodal sample corresponding to a first private entity, a multimodal sample corresponding to a first search request, and a multimodal sample corresponding to a second search request; the first private entity represents an entity having an association with a user, the first search request includes the first private entity, and the second search request does not include the first private entity; the first model is trained using the multimodal sample set to obtain a second model.

[0026] Based on the above technical solution, during the training of the first model, it relies on a certain "private" dataset for fine-tuning. This allows the trained model (i.e., the second model) to generate similar features for the same private entity in samples of different modalities, and has the ability to infer public information and visualize the features of private information; the second model can then be used to perform a visualized search, greatly improving the search effect and efficiency. As an example, a public dataset can be used to train a multimodal model to obtain a first model for generating similar features for the same non-private entity in samples of different modalities; the first model trained with the public dataset has the ability to recognize common entities.

[0027] According to the third aspect, in a first possible implementation manner of the third aspect, the use of the multimodal sample set to train the first model to obtain the second model includes: processing the multimodal samples corresponding to the first private entity and the multimodal samples corresponding to the first search request through the first model to obtain fused features; training the first model according to the fused features, the features corresponding to the first search request, and the features corresponding to the second search request to obtain the second model; wherein the features corresponding to the first search request are obtained by processing the multimodal samples corresponding to the first search request by the first model, and the features corresponding to the second search request are obtained by processing the multimodal samples corresponding to the second search request by the first model.

[0028] Based on the above technical solution, a multimodal model and private data are used to fuse private features with public features, and then align them with the corresponding multimodal features to train a public-private fusion multimodal model. In this way, the first model is fine-tuned using joint training of public and private features, which enables the trained model (i.e., the second model) to have the ability to infer concrete private information. This solves the problem that multimodal models trained only with public data sets cannot carry concrete private information.

[0029] According to a first possible implementation manner of the third aspect, in a second possible implementation manner of the third aspect, the multimodal samples corresponding to the first private entity and the multimodal samples corresponding to the first search request are processed by the first model to obtain a fused feature, including: inputting the multimodal samples corresponding to the first search request into the first model to generate a first feature; inputting the multimodal samples corresponding to the first private entity into the first model to generate a second feature; and fusing the first feature and the second feature to obtain the fused feature.

[0030] Based on the above technical solution, during the training process of the first model, the first model is used to generate private features (i.e., the second features) and public features (i.e., the first features), and the generated public features and private features are fused. The fused features can simultaneously represent public information and concrete private information, so that the first model can take into account the learning of public information and private information.

[0031] According to the first possible implementation manner of the third aspect or the second possible implementation manner of the third aspect, in the third possible implementation manner of the third aspect, the multimodal samples include samples of the first modality and samples of the second modality; the features corresponding to the first search request are obtained by processing the samples of the first modality corresponding to the first search request by the first model, and the features corresponding to the second search request are obtained by processing the samples of the first modality corresponding to the second search request by the first model; the processing of the multimodal samples corresponding to the first private entity and the multimodal samples corresponding to the first search request by the first model to obtain fused features includes: inputting the samples of the first modality corresponding to the first private entity and the samples of the second modality corresponding to the first search request into the first model to obtain the fused features.

[0032] In a fourth aspect, an embodiment of the present application provides a search device, comprising: a receiving module for receiving a first search request from a user, wherein the first search request includes a private entity; wherein the private entity represents an entity having an associated relationship with the user; a determination module for determining, among the features corresponding to at least one private entity, the features corresponding to the private entity in the first search request; wherein the features corresponding to the at least one private entity indicate data corresponding to the at least one private entity, and the data corresponding to the at least one private entity corresponds to a different modality from the first search request; a search module for obtaining search results based on the first search request and the features corresponding to the private entity in the first search request; the search results correspond to a different modality from the first search request; and a display module for displaying the search results.

[0033] According to the fourth aspect, in a first possible implementation of the fourth aspect, the search module is further used to: process the first search request and features corresponding to private entities in the first search request through a model to generate a fusion feature; the fusion feature indicates the first search request; and use data corresponding to features in the database that match the fusion feature as the search result.

[0034] According to the fourth aspect or the first possible implementation manner of the fourth aspect, in the second possible implementation manner of the fourth aspect, the device also includes: a generation module, used to obtain data corresponding to a first private entity; the first private entity is any private entity among the at least one private entity; the data corresponding to the first private entity is processed to generate features corresponding to the first private entity.

[0035] According to the second possible implementation manner of the fourth aspect, in the third possible implementation manner of the fourth aspect, the generation module is further used to: send a first prompt message to the user, where the first prompt message is used to prompt the user to select data containing the first private entity; and obtain data corresponding to the first private entity in response to the user's selection operation.

[0036] According to a second possible implementation manner of the fourth aspect, in the fourth possible implementation manner of the fourth aspect, the generation module is further used to: obtain initial data containing the first private entity; display at least one search result in response to a second search request of the user; the second search request contains the first private entity, and the at least one search result corresponds to a different modality from the second search request; the at least one search result includes the initial data; based on the user's operation of selecting the initial data in the at least one search result, use the initial data as data corresponding to the first private entity.

[0037] According to the second, third or fourth possible implementation manner of the fourth aspect, in the fifth possible implementation manner of the fourth aspect, the generation module is further used to: obtain data containing the first private entity and the second private entity; wherein the second private entity is any private entity other than the first private entity in the at least one private entity; and obtain features corresponding to the second private entity based on the data containing the first private entity and the second private entity, and features corresponding to the first private entity.

[0038] According to the foregoing various possible implementations of the fourth aspect, in a sixth possible implementation of the fourth aspect, the private entity includes: a title, a name, or a nickname associated with the user.

[0039] In the fifth aspect, an embodiment of the present application provides a search device, comprising: a first display module for displaying a first display interface; the first display interface includes a search portal; a second display module for displaying a second display interface in response to a user inputting a keyword in the search portal, the second display interface including search results corresponding to the keyword; wherein the keyword includes a private entity; the search results are obtained based on the features corresponding to the keyword and the private entity; the search results and the keyword correspond to different modalities.

[0040] According to the second aspect, in a first possible implementation of the second aspect, the device further includes: a third display module, used to display a third display interface, the third display interface including a first prompt message and at least one data; the first prompt message is used to prompt the user to select data containing a first private entity; a fourth display module, used to display a fourth display interface in response to the user's selection operation in the at least one data, the fourth display interface including first data and a first identifier, the first identifier being used to indicate that the first data is the data selected by the user.

[0041] According to the second aspect or the first possible implementation of the second aspect, in the second possible implementation of the second aspect, the second display interface also includes: second prompt information, the second prompt information is used to prompt the user to confirm whether the first search result is the search result expected by the user; the device also includes: a fifth display module, used to display a fifth display interface in response to the user's confirmation operation, the fifth display interface including the first search result and a second identifier, the second identifier is used to indicate that the first search result is the search result confirmed by the user.

[0042] In a sixth aspect, an embodiment of the present application provides a model training device, comprising: an acquisition module for acquiring a multimodal sample set and a first model; wherein the multimodal sample set comprises: a multimodal sample corresponding to a first private entity, a multimodal sample corresponding to a first search request, and a multimodal sample corresponding to a second search request; the first private entity represents an entity having an association with a user, the first search request includes the first private entity, and the second search request does not include the first private entity; a training module for training the first model using the multimodal sample set to obtain a second model.

[0043] According to the sixth aspect, in a first possible implementation manner of the sixth aspect, the training module is further used to: process the multimodal samples corresponding to the first private entity and the multimodal samples corresponding to the first search request through the first model to obtain fused features; train the first model according to the fused features, the features corresponding to the first search request, and the features corresponding to the second search request to obtain the second model; wherein the features corresponding to the first search request are obtained by processing the multimodal samples corresponding to the first search request by the first model, and the features corresponding to the second search request are obtained by processing the multimodal samples corresponding to the second search request by the first model.

[0044] According to the first possible implementation manner of the sixth aspect, in the second possible implementation manner of the sixth aspect, the training module is further used to: input the multimodal sample corresponding to the first search request into the first model to generate a first feature; input the multimodal sample corresponding to the first private entity into the first model to generate a second feature; and fuse the first feature and the second feature to obtain the fused feature.

[0045] According to the first possible implementation manner of the sixth aspect or the second possible implementation manner of the sixth aspect, in the third possible implementation manner of the sixth aspect, the multimodal samples include samples of the first modality and samples of the second modality; the features corresponding to the first search request are obtained by processing the samples of the first modality corresponding to the first search request by the first model, and the features corresponding to the second search request are obtained by processing the samples of the first modality corresponding to the second search request by the first model; the training module is further used to: input the samples of the first modality corresponding to the first private entity and the samples of the second modality corresponding to the first search request into the first model to obtain the fused features.

[0046] In the seventh aspect, an embodiment of the present application provides an electronic device, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the first aspect or one or more search methods of the first aspect, the search method of the second aspect, or the third aspect or one or more model training methods of the third aspect when executing the instructions.

[0047] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the first aspect or one or several search methods of the first aspect, the search method of the second aspect, or the second aspect or one or several model training methods of the second aspect.

[0048] In the ninth aspect, an embodiment of the present application provides a computer program product, which, when running on a computer, enables the computer to execute the above-mentioned first aspect or one or several search methods of the first aspect, the search method of the second aspect, or execute the above-mentioned second aspect or one or several model training methods of the second aspect.

[0049] The technical effects of the fourth to ninth aspects mentioned above can be found in the first, second or third aspects mentioned above. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] 1( a )-( c ) are schematic diagrams illustrating various application scenarios of a search method according to an embodiment of the present application.

[0051] FIG2 shows a flow chart of a search method according to an embodiment of the present application.

[0052] FIG3 shows a flow chart of a method for obtaining search results according to an embodiment of the present application.

[0053] 4( a )-( c ) are schematic diagrams illustrating a search method according to an embodiment of the present application.

[0054] FIG5 shows a flow chart of a method for constructing a private knowledge set according to an embodiment of the present application.

[0055] FIG6 shows a flowchart of an explicit collection method according to an embodiment of the present application.

[0056] 7( a )-( c ) are schematic diagrams illustrating an explicit data collection method according to an embodiment of the present application.

[0057] 8( a )-( c ) are schematic diagrams illustrating an explicit data collection method according to an embodiment of the present application.

[0058] FIG9 shows a flow chart of an implicit collection method according to an embodiment of the present application.

[0059] 10( a )-( c ) are schematic diagrams illustrating an implicit acquisition method according to an embodiment of the present application.

[0060] FIG11 shows a flowchart of a model training method according to an embodiment of the present application.

[0061] FIG12 shows a flowchart of a model training method according to an embodiment of the present application.

[0062] FIG13 shows a flow chart of a model training method according to an embodiment of the present application.

[0063] FIG14 shows a structural diagram of a search device according to an embodiment of the present application.

[0064] FIG15 shows a structural diagram of a model training device according to an embodiment of the present application.

[0065] FIG16 shows a schematic structural diagram of an electronic device 100 according to an embodiment of the present application.

[0066] FIG17 shows a software structure block diagram of the electronic device 100 according to an embodiment of the present application.

[0067] FIG18 shows a schematic structural diagram of another electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0068] Various exemplary embodiments, features, and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0069] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0070] In this application, "at least one" means one or more, and "more" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: including the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0071] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0072] The word "exemplary" is used herein to mean "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior or preferred over other embodiments. Furthermore, numerous specific details are provided in the following detailed description to better illustrate the present application. Those skilled in the art will appreciate that the present application may be practiced without some of these specific details.

[0073] In order to better understand the solutions of the embodiments of the present application, the relevant terms and concepts that may be involved in the embodiments of the present application are first introduced below.

[0074] 1. Multimodal Data

[0075] Multimodal data refers to data obtained from different fields or perspectives for the same descriptive object. Each field, perspective, form of existence, or information source describing this data is called a modality. Data composed of two or more modalities is called multimodal data. For example, modal data can include text, images, video, audio, and other data.

[0076] 2. Multimodal Search

[0077] Also known as multimodal search or multimodal retrieval; it is a technology that searches by using query statements (such as keywords) that are inconsistent with the modality of the data to be searched, for example, searching for images with text.

[0078] 3. Multimodal search within the terminal

[0079] A technology that enables multimodal search within terminals such as mobile phones, tablets, and personal computers (PCs).

[0080] 4. Cross-end multimodal search

[0081] A technology that distributes search requests to multiple terminals through a data protocol, enabling multimodal search and result aggregation across terminals.

[0082] 5. End-to-end multimodal search

[0083] A technology that combines a terminal with a remote server such as a cloud server through a certain data protocol, thereby realizing multimodal joint search through the terminal and the cloud server.

[0084] First, the application scenarios to which the search method in the embodiment of the present application can be applied are exemplarily described below. Figures 1(a)-(c) are schematic diagrams showing various application scenarios of the search method according to an embodiment of the present application.

[0085] Scenario 1: In-terminal multimodal search scenario, in which a user can perform local gallery searches, global searches, and the like on a terminal device that stores the user's private photos. For example, the terminal device may be a personal computer, laptop, smartphone, tablet, IoT device, or portable wearable device. IoT devices may include smart speakers, smart TVs, smart air conditioners, and smart car-mounted devices, and portable wearable devices may include smart watches, smart bracelets, and head-mounted devices. As an example, taking in-terminal text-based image search as an example, as shown in FIG1( a ), mobile phone 101 may display a user search interface, which includes an entry for image search requests. The user may enter keywords through the entry for the image search request, thereby triggering mobile phone 101 to search the local gallery for the user's desired images and display the searched images on mobile phone 101.

[0086] Scenario 2: Cross-end multimodal search scenario. In this scenario, a user can use one terminal device to search the gallery on other terminal devices, perform global searches, etc., wherein a connection is established between the terminal device and the other terminal device, and the other terminal device stores the user's private photos. As an example, taking cross-end text-to-image search as an example, as shown in Figure 1(b), a connection is established between mobile phone 101 and personal computer 102. Mobile phone 101 can display a user search interface, which includes an entry for image search requests. The user can enter keywords through the entry for the image search request, and mobile phone 101 transmits the keywords to personal computer 102, thereby triggering personal computer 102 to search the personal computer gallery for the user's desired images, and the searched images can be displayed on mobile phone 101.

[0087] Scenario 3: End-to-end multimodal search scenario. In this scenario, a user can use a terminal device to perform a gallery search, a global search, etc. on the gallery on the cloud server. A connection is established between the terminal device and the cloud server, and the cloud server stores the user's private photos. For example, the cloud server can be an independent server or a server cluster composed of multiple servers. As an example, taking the end-to-end text-to-image search as an example, as shown in Figure 1(c), the mobile phone 101 can establish a connection with the cloud server 103 through the network. The mobile phone 101 can display a user search interface, which includes an entry for image search requests. The user can enter a keyword through the entry for the image search request, and the mobile phone 101 transmits the keyword to the cloud server 103, thereby triggering the cloud server 103 to search the cloud gallery for the image the user wants, and the searched image can be displayed on the mobile phone 101.

[0088] Taking the in-terminal text-to-image search scenario shown in FIG1( a ) as an example, in the related art, there are mainly two in-terminal text-to-image search methods.

[0089] Method 1 uses fixed tags for on-device multimodal search. In this method, for images in the gallery, inherent tags (such as time and location) are combined with tags inferred by models such as object detection, image classification, and optical character recognition (OCR) to form a closed tag set. An inverted index is then established based on this closed tag set. The query tag is then calculated. Furthermore, using text segmentation technology, the keywords entered by the user are decomposed into multiple tags. Finally, an inverted tag search is performed based on the multiple tags obtained above to find images that meet the requirements and return them to the user.

[0090] This method can realize the image search function, but it still has the following shortcomings: (1) It only supports searches based on closed tag sets. When users search, keywords need to be strictly input according to the tags and strictly matched. It has poor flexibility and is difficult to solve matching scenarios such as ambiguity, fuzziness, and synonyms, and cannot achieve semantic matching; (2) The tags in the closed tag set are very limited and cannot effectively represent the semantics of the image. It cannot meet the diversified search requests of users in most cases. Even if traditional natural language understanding (NLU) is used as keywords, it still needs to be customized for different search businesses, lacks versatility, and has poor scalability; (3) If the closed tag set is expanded, the inference model needs to be replaced accordingly, which is difficult to maintain.

[0091] Method 2 uses an open, on-device multimodal search method. In this method, a multimodal model is first trained using publicly available image-text pair data to obtain a multimodal model with the same semantic representation of image-text pairs in a high-dimensional space. Then, the images in the image library are inferred using the multimodal model to obtain corresponding high-dimensional spatial feature vectors, forming the image base library data. Furthermore, the multimodal model is used to infer the high-dimensional spatial feature vectors of user keywords. Finally, the similarity between the generated high-dimensional spatial feature vectors and the image feature vectors in the image semantic base library is calculated to obtain the image with the highest similarity, thus completing the image library search.

[0092] This approach redefines image search technology from a semantic perspective, providing a more comprehensive representation of images and meeting users' daily open-ended search needs (i.e., allowing for arbitrary input without the need for specific tags). However, it still suffers from several drawbacks: a lack of concrete search capabilities, meaning it cannot match entities in the query text with specific images in reality, making it difficult to achieve a true semantic understanding of the content and user requests. For example, if a user enters the keyword "my wife and I playing badminton," a search will display photos of two people playing badminton, but there's no guarantee that these two people are the user and their wife. For example, the search results may include photos of two people playing badminton, such as "my wife and I playing badminton" and "my daughter and I playing badminton," among others. The user still needs to further select photos from the numerous search results to obtain the desired "my wife and I playing badminton" photos, resulting in low search efficiency, low accuracy, and a poor user experience.

[0093] In order to solve the above-mentioned problems in the related art, the embodiment of the present application provides a search method (see below for detailed description). This can effectively solve the problem that users cannot perform figurative searches when searching through terminal devices, and improve the user's search experience. Compared with the terminal-side search method with fixed tags, this search method does not require a closed tag set. Users can enter keywords arbitrarily when searching, which is highly flexible; it can meet the user's diversified search requests and has higher scalability; this search method goes beyond the scope of tags and does not require continuous enrichment and exhaustive enumeration of tags. For different terminal devices, there is no need to replace the model, and maintenance is easier. Compared with the open-end multimodal search method, this search method can correspond the entities in the query text to the specific images in reality, realize figurative search, and has high search efficiency, high accuracy, and better user experience.

[0094] It should be noted that the above-mentioned application scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Ordinary technicians in this field can know that for the emergence of other similar or new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems. For example, they can be used to search for private data authorized by users in terminal devices.

[0095] The following will introduce the technical solutions provided by this application from two aspects: the application side and the training side. First, the search method provided by the embodiment of this application will be described in detail from the application side.

[0096] FIG2 is a flowchart of a search method according to an embodiment of the present application. Exemplarily, the method may be executed in a terminal device, for example, the mobile phone 101 in FIG1 . As shown in FIG2 , the method may include the following steps:

[0097] Step S201: Receive a first search request from a user, where the first search request includes a private entity.

[0098] Exemplarily, the first search request may be data in a text modality (such as keywords such as text, words, sentences, paragraphs consisting of sentences), an image modality (such as an image containing text), a video modality (such as a video containing text), or an audio modality (such as voice).

[0099] For example, the first search request may include one or more private entities and one or more non-private entities. A private entity represents an entity associated with the user, such as "I," "wife," "husband," "uncle," "aunt," "son," or "mom." Another example may be a name or nickname associated with the user, such as "Zhang San" or "Dian Dian." It is understood that for different users, the same private entity corresponds to different real-world images, such as different people or objects. For example, for user A, the private entity "I" corresponds to user A, while for user B, the private entity "I" corresponds to user B. A non-private entity represents an entity not directly associated with the user, such as general objects or locations, such as "shoes," "badminton," "swimming," and "Shanghai." For different users, the same non-private entity corresponds to the same real-world image, such as the same person or object.

[0100] As an example, a user can send a first search request in text mode to a terminal device, and the first search request in text mode may include a private entity; for example, as shown in FIG1(a), a user can enter a keyword through the search bar in the gallery of mobile phone 101, and the keyword may include a private entity.

[0101] As another example, a user may send a first search request in audio mode to a terminal device, and the terminal device may perform a modal transformation on the first search request in audio mode to obtain a first search request in text mode, where the first search request in text mode includes private entities. For example, a user may input a search voice through the search entry on the negative first screen of mobile phone 101, and mobile phone 101 may process the search voice input by the user and convert it into a keyword, where the keyword includes a private entity.

[0102] For example, after the user sends a first search request to the terminal device, the terminal device can parse the first search request to obtain at least one entity, and then match the at least one entity with at least one private entity in the terminal device. The one or more entities that are successfully matched are the private entities in the first search request.

[0103] For example, taking text-to-image search as an example, at least one private entity in the terminal device may include: "I", "wife", "son", "uncle", "mother", "Zhang San", etc. The first search request may be the keyword "my photos". The terminal device parses the keyword "my photos" input by the user, and can extract the two entities "I" and "photos" in the keyword "my photos", and then match "I" and "photos" with the above-mentioned multiple private entities respectively, among which "I" is matched successfully, that is, the private entity in the keyword "my photos" is "I".

[0104] Step S202: Determine, among the features corresponding to at least one private entity, the features corresponding to the private entity in the first search request; wherein the features corresponding to the at least one private entity indicate data corresponding to the at least one private entity, and the data corresponding to the at least one private entity corresponds to a different modality than the first search request.

[0105] Exemplarily, a private knowledge set may be stored in the terminal device, which may include at least one private entity and features (also called feature vectors) corresponding to each private entity in the at least one private entity; the terminal device selects the features corresponding to the private entity in the private knowledge set based on the private entity in the first search request received from the user.

[0106] Exemplarily, the features corresponding to at least one private entity in the private knowledge set are obtained by processing data corresponding to the at least one private entity by a model in a terminal device. For example, the first search request may be data in a textual modality, and the data corresponding to the at least one private entity may be data in an image modality. Another example is that the first search request may be data in an audio modality, and the data corresponding to the at least one private entity may be data in an image modality. It is understood that for data corresponding to a private entity in a particular modality, the features corresponding to the private entity obtained by the model reflect the characteristics of the private entity in that modality. For example, for a selfie of me, the features corresponding to the private entity "I" obtained by the model reflect the visual characteristics of "I" in the image, such as color, shape, position, or size. The training process of the model in the terminal device is described below. Exemplarily, the trained model can be ported to the terminal device. For example, if the first search request is data in a textual modality, image modality data containing the private entity in the terminal device can be input into the model to obtain the features corresponding to the private entity. Associations between the private entity and the features corresponding to the private entity are then established, thereby obtaining the private knowledge set. For example, we can input photos 1 through 10 from the local gallery on a phone into the model. Each of these photos contains at least one private entity. It's understood that a photo can contain one or more private entities. For example, a photo of me and my wife might contain two private entities, "me" and "wife." This allows us to obtain the features corresponding to each private entity in photos 1 through 10. We can then determine the corresponding relationships between each private entity and its corresponding features, thereby establishing a private knowledge set.

[0107] As an example, at least one private entity in the private knowledge set may include "I", "wife", "son", "uncle", "mother", and "Zhang San"; the private knowledge set may include characteristics of "I", characteristics of "wife", characteristics of "son", characteristics of "uncle", characteristics of "mother", and characteristics of "Zhang San". If the private entity in the first search request received by the terminal device in the above steps is "I", the characteristics of "I" can be selected in the private knowledge set.

[0108] Step S203: Obtain search results based on the first search request and the features corresponding to the private entities in the first search request.

[0109] Exemplarily, the terminal device may process the first search request and the features corresponding to the private entity in the first search request respectively through the above model, and fuse the processing results, thereby obtaining search results based on the fused processing results.

[0110] The search results and the first search request correspond to different modalities. For example, the search results and the data corresponding to the at least one private entity may correspond to the same modality. As an example, the first search request may be data in a text modality, and the search results may be data in an image modality. For example, the first search request may be the keyword "my photos," and the search results may be one or more photos.

[0111] FIG3 shows a flow chart of a method for obtaining search results according to an embodiment of the present application. As shown in FIG3 , step S203 may include the following steps:

[0112] Step S20301: Process the first search request and features corresponding to private entities in the first search request through a model to generate a fused feature; the fused feature indicates the first search request.

[0113] Exemplarily, the model can be pre-configured in the terminal device, and the model can generate similar features for the same private entity in data of different modalities. Exemplarily, the terminal device can input the first search request into the model to generate a first feature, and then fuse the first feature with the feature corresponding to the private entity in the first search request to generate a fused feature. As an example, the model may include: a sub-model corresponding to the first modality and a sub-model corresponding to the second modality; wherein the first modality is the modality corresponding to the data corresponding to the private entity in the first search request, such as an image modality, and the second modality is the modality corresponding to the first search request, such as a text modality. The terminal device can input the feature corresponding to the first search request into the sub-model corresponding to the second modality to obtain the first feature, and then fuse the first feature with the feature corresponding to the private entity in the first search request to generate a fused feature.

[0114] It is understandable that for different users, since the same private entity may correspond to different images in reality, that is, the features corresponding to the same private entity may be different, that is, the features corresponding to the private entity can represent concrete private information, which is related to the specific user and can also be called private features. The first feature reflects the semantics of the text itself in the first search request; it is understandable that for different users, if the first search request is the same, then the semantics of the text itself in the first search request is the same, and the first feature generated by inputting the first search request into the above model is the same. Therefore, the first feature can represent public information, which is not related to the specific user and can also be called a public feature; furthermore, the private feature and the public feature can be fused to obtain a fused feature, which can also be called a public-private fusion feature.

[0115] For example, existing fusion methods such as adapter, summation, and attention can be used to fuse the first feature and the feature corresponding to the private entity in the first search request to obtain a fused feature. This fused feature is a feature in a high-dimensional space that can represent both private and public features, thereby achieving complementarity between private and public features.

[0116] Step S20302: Use the data corresponding to the features matching the fusion features in the database as search results.

[0117] Exemplarily, the database may include a plurality of data and features corresponding to each of the plurality of data; wherein the plurality of data corresponds to a different modality than the first search request.

[0118] Exemplarily, the terminal device may calculate the similarity between the fused feature and the features corresponding to each data in the database, and use the feature in the database with the highest similarity to the fused feature or the top K similarity ranking features as the features that match the fused feature. For example, the features may be sorted from high to low by similarity, so that the features corresponding to the top K similarities are used as the features that match the fused feature; the data corresponding to the features that match the fused feature is the search result. It is understandable that the number of features that match the fused feature may be one or more, wherein different features that match the fused feature correspond to different search results, i.e., the number of search results may be one or more.

[0119] As an example, the database can be a picture gallery in a terminal device, including multiple pictures and features corresponding to each of the multiple pictures. The above-mentioned fusion features can be used to calculate the similarity with the features in the database, and the features with the highest similarity in the database or the top K features in similarity ranking are used as features matching the fusion features, and then the pictures corresponding to the features matching the fusion features are used as searched pictures.

[0120] Exemplarily, the aforementioned model in the terminal device can be used to pre-infer data in different modalities corresponding to the first search request in the terminal device, thereby obtaining features corresponding to each data item. The features corresponding to each data item can be high-dimensional spatial features. A correspondence between each data item and its corresponding features is established, and based on this correspondence, each data item and its corresponding features are stored in a database of the terminal device. Exemplarily, the terminal device can input image modality data into the sub-model corresponding to the image modality in the model, thereby generating features corresponding to the data in that image modality.

[0121] For example, the terminal device can trigger operations such as inferring the features corresponding to each data according to upper-layer scheduling at an appropriate time, for example, when the computing power of the terminal device permits or during a time period when the user does not use the terminal device, thereby updating the database.

[0122] As an example, the database can be a gallery in a mobile phone. It is understandable that there are usually new pictures in the gallery every day. When the computing power of the mobile phone allows, or when the user is not using the mobile phone, the above model can be used to infer the features corresponding to the new pictures, thereby reducing the impact on the user's normal use of the mobile phone; for example, when the user is not using the mobile phone in the early morning, the sub-model corresponding to the image modality in the above model can be used to infer the pictures in the gallery or the new pictures in sequence, obtain the features corresponding to each picture, and store them in the terminal device.

[0123] In this way, through the above steps S20301-S20302, the terminal device uses the model to infer the first search request and generate a fusion feature, which can simultaneously represent public information and concrete private information; as an example, the first search request can be processed by the model to generate a public feature (i.e., the first feature), and then the generated public feature and the private feature (i.e., the feature corresponding to the private entity in the first search request) are fused to generate a fusion feature; thereby, for different users, the search results that the user wants can be searched, and a concrete search based on the fusion of private features and public features can be realized.

[0124] Illustratively, when the feature corresponding to at least one private entity in step S202 does not include the feature corresponding to the private entity in the first search request, the first search request may be processed by the model to obtain search results.

[0125] Step S204: Display search results.

[0126] Exemplarily, the terminal device may display the search results. For example, if the search results are image-modal data, i.e., if the search results are images, the images may be displayed to the user via the display screen on the terminal device. Exemplarily, the images may be displayed directly, or a thumbnail of the images may be displayed, etc., without limitation.

[0127] In this way, through the above steps S201-S204, the features corresponding to the private entity in the user's search request are determined from the features corresponding to the at least one private entity, and search results are obtained based on the features corresponding to the private entity and the first search request. Since the features corresponding to the at least one private entity indicate the data corresponding to the at least one private entity, the data corresponding to the at least one private entity can correspond to the same modality as the search results, so that the private entity in the search request can be matched with the specific image in reality (i.e., the specific image of the private entity in the search results), thereby achieving a concrete search; solving the problem that the terminal device cannot perform a concrete search during the search process, improving the search effect and search efficiency, and enhancing the user's search experience. As an example, the model can be used to infer the private entity in the first search request and the first search request to obtain public-private fusion features, and then the public-private fusion features can be used to determine the search results in the database. In addition, there is no need for a closed tag set, which is highly flexible; it can meet the user's diverse search requests and has higher scalability; for different terminal devices, there is no need to replace the model, which is easier to maintain.

[0128] The present application also provides another search method, which may include: a terminal device may display a first display interface; the first display interface includes a search entry; further, in response to a user inputting a keyword into the search entry, the terminal device displays a second display interface, the second display interface including a first search result corresponding to the keyword; wherein the keyword includes a private entity; the first search result is obtained based on the keyword and the features corresponding to the private entity in the keyword; the first search result corresponds to a different modality than the keyword. The number of first search results may be one or more; illustratively, the process of obtaining the first search result may refer to the relevant description in step S203 above.

[0129] For example, taking the mobile phone 101 in Figure 1(a) as the terminal device, Figures 4(a)-(c) show schematic diagrams of a search method according to an embodiment of the present application. As shown in Figure 4(a), the user opens the local gallery in the mobile phone 101 and displays the first display interface, wherein a search box (i.e., a search entry) is provided above the first display interface; the first display interface can also display part or all of the images in the local gallery; as shown in Figure 4(b), the user can enter the keyword "I play badminton" in the search box in the local gallery and click the search icon in the search box to trigger the mobile phone 101 to perform a search operation; illustratively, the mobile phone 101 parses the keyword "I play badminton" and obtains the private entity "I" in the keyword "I play badminton" among multiple private entities. Then, the mobile phone 101 can search for the private entity "I" in the features corresponding to the multiple private entities. , determine the features corresponding to the private entity "I", and can search for pictures of "I playing badminton" in the gallery according to the keyword "I play badminton" and the features corresponding to the private entity "I". For example, the mobile phone 101 can fuse the public features obtained by inputting the keyword "I play badminton" into the model with the private features corresponding to the private entity "I" to obtain a fused feature, and then determine the picture corresponding to the feature matching the fused feature in the gallery as the picture of "I playing badminton" that the user wants; as shown in Figure 4(c), the mobile phone 101 can then display a second display interface, and the searched picture of "I play badminton" can be displayed in the second display interface, thereby completing the concrete search.

[0130] The following describes in detail possible implementation methods for establishing the private knowledge set.

[0131] For example, the number of private entities contained in the private knowledge set can be set as needed. For example, in view of the limited number of images in the gallery of the terminal device, in order to improve the effectiveness and efficiency of data collection, the types of private entities contained in the images in the gallery of most users can be statistically obtained based on the minimization principle (for example, "I", "wife", "husband", "uncle", "aunt", "son", "mother" and other common characters related to users), so as to determine the number of private entities in the minimum private knowledge set, and then generate the corresponding features of each private entity to complete the construction of the minimum private knowledge set. This minimum private knowledge set can solve the vast majority of the concrete search requirements in the field of image search, thereby meeting the search needs of different users. For example, the private entities contained in the private knowledge set can also be added or reduced according to needs, and there is no limitation on this.

[0132] FIG5 shows a flow chart of a method for constructing a private knowledge set according to an embodiment of the present application. The method can be executed on a terminal device, for example, the mobile phone 101 in FIG1 . As shown in FIG5 , the method may include the following steps:

[0133] Step S501: Acquire data corresponding to a first private entity.

[0134] The first private entity is any one of the at least one private entity described above; the data corresponding to the first private entity corresponds to a different modality than the first search request. The data corresponding to the first private entity may include only the first private entity or may include multiple private entities including the first private entity, without limitation.

[0135] It is understandable that the data corresponding to the same private entity may be different for different users. That is, for a particular user, the data corresponding to the private entity may serve as concrete private information related to that user. For example, taking the first private entity as "son," for user A, the data corresponding to "son" obtained by user A's terminal device may be a photo of user A's son; for user B, the data corresponding to "son" obtained by user B's terminal device may be a photo of user B's son.

[0136] Exemplarily, the data corresponding to the first private entity may be acquired in the following collection manner.

[0137] Method 1: Data corresponding to the first private entity can be obtained through explicit collection. For example, the terminal device can guide the user to make a selection by providing a prompt message or other user-perceivable means, and then obtain the data corresponding to the first private entity based on the user's selection. This allows the data corresponding to the first private entity to be obtained within the scope permitted by privacy and security.

[0138] Method 2: The data corresponding to the first private entity can be obtained by implicit collection. For example, the terminal device can obtain the data corresponding to the first private entity by inferring data containing common sense information (for example, wedding photos, parent-child photos, selfies, etc.), thereby obtaining the data corresponding to the first private entity without the user's perception; considering that the confidence level of the obtained data corresponding to the first private entity may be low, the terminal device can further respond to the search request actively triggered by the user and display the search results to the user. Based on the user's selection operation in the search results, it can be determined whether the data corresponding to the first private entity with a lower confidence level is data confirmed by the user, and the data corresponding to the first private entity with a higher confidence level can be obtained; thereby, the data corresponding to the first private entity can be obtained within the scope permitted by privacy security.

[0139] Step S502: Process the data corresponding to the first private entity to generate features corresponding to the first private entity.

[0140] For example, the terminal device can process the data corresponding to the first private entity using a model in the terminal device to generate features corresponding to the first private entity. It is understood that for data corresponding to the first private entity in a certain modality, the features corresponding to the first private entity generated by the model reflect the characteristics of the first private entity in that modality. For example, for a photo like "My Selfie," when the photo is input into the model, the features corresponding to "I" generated reflect the characteristics of "I" in the image.

[0141] Exemplarily, the model may include sub-models corresponding to multiple modalities, and the terminal device may input data corresponding to the first private entity of a certain modality into the sub-model corresponding to the modality; for example, the photo "my selfie" may be input into the sub-model corresponding to the image modality, thereby generating features corresponding to the image modality "I".

[0142] It can be understood that since the data corresponding to the same private entity contains concrete private information related to the user; the features corresponding to the generated first private entity can be used as private knowledge related to the user. In the multimodal search process, this private knowledge can be used as known information to realize concrete search. For example, the private knowledge and search request can be input into the model to realize concrete search.

[0143] In some examples, when the data corresponding to the first private entity includes multiple private entities including the first private entity, the terminal device may further split the data into multiple data containing only one private entity, and then input each data containing only one private entity into the model to obtain features corresponding to each private entity. For example, the terminal device may identify the age of the person in the photo, thereby splitting "a photo of me and my son" into a photo containing only me and a photo containing only my son, and inputting these two photos into the model respectively to obtain features corresponding to the private entity "me" and features corresponding to the private entity "son" respectively; alternatively, the terminal device may further input the data corresponding to the first private entity into the model to obtain features corresponding to the multiple private entities. For example, the terminal device may input "a photo of me and my son" into the model to obtain features corresponding to the two private entities "me" and "son".

[0144] In this way, through the above steps S501-S502, the data corresponding to the first private entity is obtained. As an example, the data corresponding to the first private entity can be obtained by explicit collection or implicit collection, so as to collect concrete private information within the scope permitted by privacy security; and the data corresponding to the first private entity is processed to generate features corresponding to the first private entity, thereby obtaining features corresponding to the first private entity; exemplarily, the features corresponding to the first private entity can be stored in the terminal device; thereby providing support for quickly determining the features corresponding to the private entity in the subsequent process of the terminal device performing multimodal search.

[0145] Furthermore, the terminal device can determine the features corresponding to other private entities in the private knowledge set based on the features corresponding to the generated private entity, thereby automatically expanding the private knowledge set and completing the construction of the private knowledge set.

[0146] As an example, for any private entity in the private knowledge set, the terminal device can repeatedly execute the above steps S501-S502, obtain the data corresponding to each private entity in turn, and use the model to generate the features corresponding to each private entity, thereby traversing each private entity, expanding private knowledge, and completing the construction of the private knowledge set.

[0147] As another example, the terminal device can also further obtain features corresponding to other private entities in the private knowledge set by executing the following steps S503 and S504 on the basis of generating features corresponding to the first private entity, thereby expanding the private knowledge set and finally obtaining features corresponding to each private entity to complete the construction of the private knowledge set.

[0148] Step S503: Acquire data including the first private entity and the second private entity.

[0149] The second private entity is any private entity among the at least one private entity except the first private entity.

[0150] In this step, the terminal device can automatically obtain data that includes both the first private entity and the second private entity within the scope of privacy security authorized by the user. It can be understood that, for a certain user, compared with data that only includes the first private entity or data that only includes the second private entity, the data that includes both the first private entity and the second private entity contains more concrete private information related to the user. For example, the data that includes the first private entity and the second private entity can be photos in the gallery that include multiple private entities, such as "wedding photos", "family photos" or "group photos". Exemplarily, the terminal device can automatically determine the "family photo" in the gallery in combination with common sense information. The photo can include multiple private entities such as "me", "wife", "son" or "daughter".

[0151] Step S504: Obtain features corresponding to the second private entity based on the data including the first private entity and the second private entity and the features corresponding to the first private entity.

[0152] Exemplarily, the terminal device can split the data containing the first private entity and the second private entity into two independent data, namely, data containing only the first private entity and data containing only the second private entity, and input these two independent data into the model respectively to obtain the features corresponding to the two private entities, namely, the features corresponding to the first private entity and the features corresponding to the second private entity; and then, based on the known features corresponding to the first private entity (i.e., the known private knowledge), the features corresponding to the second private entity can be determined from the features corresponding to the two private entities, thereby realizing the automatic expansion of the private knowledge set.

[0153] As an example, a terminal device can obtain a "selfie" from a gallery, that is, an image corresponding to the private entity "I", and input the "selfie" into the model to generate features corresponding to the image modality "I". Furthermore, the terminal device can also obtain a "wedding photo" from the gallery, that is, an image containing the private entities "I" and "wife", and extract features corresponding to the two private entities in the "wedding photo", that is, features corresponding to the image modalities "I" and "wife". Finally, the terminal device can combine the features corresponding to the image modality "I" generated by the "selfie" above to obtain features corresponding to the image modality "wife".

[0154] In this way, the terminal device can collect data corresponding to private entities (i.e., concrete private information) within the scope permitted by privacy and security through the above steps S501-step S502, and then generate private knowledge; for example, the private information can be processed by a model to generate private knowledge; and then the terminal device can also automatically expand the generated private knowledge set; as an example, the terminal device can automatically expand the private knowledge set based on the acquired data containing multiple private entities and the generated private knowledge through common sense and models through the above steps S503-step S504, so as to improve the concrete private knowledge within the scope permitted by privacy and security with as little collection as possible and using a small amount of generated private knowledge, thereby achieving effective use of data in situations where user feedback data is scarce, and providing support for concrete search in the process of terminal devices performing multimodal search. In some examples, after executing the above steps, the terminal device can further determine whether to traverse each private entity in the minimum private knowledge set. If all private entities are traversed, that is, the features corresponding to all private entities are generated, then the construction of the minimum private knowledge set is completed. Otherwise, repeat the above steps S503-step S504 until all private entities are traversed and the constructed minimum private knowledge set is stored in the terminal device; illustratively, a triple <1#, 2#, character relationship> storage method can be adopted, where 1# represents the features corresponding to private entity 1 in the minimum private knowledge set, 2# represents the features corresponding to private entity 2 in the minimum private knowledge set, and the character relationship represents the relationship between private entities 1 and 2. 1# and 2# are stored in the database in the form of key-value pairs.

[0155] The following specifically describes the process of acquiring data corresponding to the first private entity in an explicit acquisition manner when constructing the private knowledge set.

[0156] FIG6 shows a flow chart of an explicit data collection method according to an embodiment of the present application. The method can be executed on a terminal device, for example, on the mobile phone 101 in FIG1 . As shown in FIG6 , the method can include the following steps:

[0157] Step S601: Send a first prompt message to the user.

[0158] The first prompt message is used to prompt the user to select data containing the first private entity. For example, the first prompt message can be in the form of voice, text, vibration, or video, etc., which is not limited to this. It is understood that the first prompt message needs to prompt the user within the scope of privacy and security authorized by the user.

[0159] Exemplarily, the terminal device may send a first prompt message to the user when the user searches for pictures in the gallery for the first time, prompting the user to select data containing the first private entity in the gallery; or, the terminal device may also send a first prompt message to the user when it detects that the user usually spends a long time searching for pictures in the gallery, prompting the user to select data containing the first private entity; or, the terminal device may also send a first prompt message to the user when the number of pictures in the gallery exceeds a certain number.

[0160] Exemplarily, the terminal device may select one or more pieces of prompt information in the prompt information set as the first prompt information based on the prompt information set.

[0161] Step S602: In response to a selection operation by the user, data corresponding to the first private entity is obtained.

[0162] Exemplarily, after the terminal device sends the first prompt information to the user, the user can select corresponding data or corresponding prompt options according to the first prompt information, and the terminal device uses the data selected by the user as the data corresponding to the first private entity.

[0163] In one possible implementation, the terminal device displays a third display interface, the third display interface including first prompt information and at least one data; the first prompt information is used to prompt the user to select data containing the first private entity; in response to the user selecting the at least one data, a fourth display interface is displayed, the fourth display interface including the first data and a first identifier, the first identifier being used to indicate that the first data is the data selected by the user. Exemplarily, the identifier can be text, color, brightness, graphics, size, or the like. For example, if the first data is an image, it can be identified by highlighting it, or by adding a border around the image.

[0164] As an example, the terminal device may send prompt information such as "Please select your photo", "Please select a photo of your son", "My wife and I play badminton", "Please select a wedding photo", or "Please select a family photo" to the user to prompt the user to select the corresponding picture in the gallery. If the user selects one or more pictures, the terminal device uses the one or more pictures as the data of the image modality corresponding to the private entity contained in the prompt information. For example, taking the terminal device as the mobile phone 101 in Figure 1(a) above, and the first private entity being "I", Figures 7(a)-(c) show a schematic diagram of an explicit collection method according to an embodiment of the present application; as shown in Figure 7(a), the user enters the local gallery of the mobile phone 101 and displays a first display interface, with a search box (i.e., a search entry) provided above the first display interface; the first display interface can also display some or all images in the local gallery, and the user can click on the search box in the first display interface to enter keywords in the search box; as shown in Figure 7(b), the mobile phone 101 can detect that the user clicks on the local gallery for the first time. When the search box is selected in the map library, a third display interface is displayed, which displays some or all images in the local library, and may also display a prompt message "Please select your photo" to guide the user to select his or her own photo from the displayed images; the user can select his or her own photo according to the prompt message, and in response to the user's selection operation, a fourth display interface is displayed, as shown in Figure 7(c), the user selects picture 2; the mobile phone 101 can determine that picture 2 is a photo of the user himself or herself, and add a border around picture 2 to indicate that picture 2 is the picture selected by the user, thereby obtaining the image modal data corresponding to the private entity "me".

[0165] In one possible implementation, the terminal device may display a first display interface; the first display interface includes a search entry; further, in response to the user inputting a keyword in the search entry, the terminal device displays a second display interface, the second display interface including a first search result corresponding to the keyword and a second prompt message; wherein the keyword includes a private entity; the first search result is obtained based on the keyword and the features corresponding to the private entity in the keyword; the first search result and the keyword correspond to different modalities; the second prompt message is used to prompt the user to confirm whether the first search result is the search result expected by the user. In response to the user's confirmation operation, the terminal device displays a fifth display interface, the fifth display interface including the first search result and a second identifier, the second identifier being used to indicate that the first search result is the search result confirmed by the user. Exemplarily, the process of obtaining the first search result may refer to the relevant description in the above step S203.

[0166] As an example, the terminal device can detect that the user actively searches for pictures in the gallery and issue a prompt message. For example, the user enters keywords such as "daughter", "photo of son" or "my wife and I playing badminton". The terminal device responds to the user's query operation, searches in the gallery, displays the searched photos, and issues a prompt message "Is this the photo you want to search for?" and displays the prompt options "Yes" and "No". If the user selects "Yes", the terminal device uses the photo as the image modal data corresponding to the private entity "son". For example, taking the terminal device as the mobile phone 101 in Figure 1(a) above, and the first private entity being "son", Figures 8(a)-(c) show schematic diagrams of an explicit collection method according to an embodiment of the present application; as shown in Figure 8(a), the user enters the local gallery of the mobile phone 101 and displays a first display interface. A search box (i.e., a search entry) is provided above the first display interface, and the keyword "son" can be entered in the search box of the first display interface; as shown in Figure 8(b), the mobile phone 101 performs a search operation and displays the searched picture 1, the prompt statement "Is this the photo you want to search for?" and the prompt options "Yes" and "No" on the screen; in response to the user clicking the option "Yes", the fifth display interface is displayed, as shown in Figure 8(c), then the mobile phone 101 can determine that the picture 1 is a photo of the user's son, and add a border around the picture 1 to indicate that the picture 1 is the picture confirmed by the user, thereby obtaining the image modality data corresponding to the private entity "son".

[0167] As another example, a button or entrance can be set in the gallery of the terminal device. If the user finds that the image search effect is not good, the button or entrance can be used to trigger the terminal device to issue a first prompt message, such as "Please select photos of Zhang San studying"; if the user selects one or more pictures, the terminal device will use the one or more pictures as the image modal data corresponding to the private entity "Zhang San".

[0168] In this way, through the above steps S601 and S602, the data corresponding to the first private entity is obtained in an explicit collection method that can be perceived by the user. Since the private information is subjectively selected by the user, the private information collected in this way has the characteristic of high confidence; thereby achieving the collection of high-confidence concrete private information through display guidance within the scope of user privacy security.

[0169] Furthermore, the terminal device may also execute step S603, input data corresponding to the first private entity into the model, and generate features corresponding to the first private entity; thereby generating private knowledge from the concrete private information obtained by the above explicit collection method.

[0170] Possible implementations of step S603 may refer to the relevant description in the above step S502.

[0171] Furthermore, the terminal device may also obtain features corresponding to other private entities in the private knowledge set by executing the following steps S604-S605, and automatically expand the private knowledge set, thereby completing the construction of the private knowledge set.

[0172] Step S604: Acquire data including the first private entity and the second private entity.

[0173] Step S605: Obtain features corresponding to the second private entity based on the data including the first private entity and the second private entity and the features corresponding to the first private entity.

[0174] The specific implementation process of steps S604-S605 can refer to the relevant descriptions in steps S503-S504 in Figure 5 above.

[0175] In this way, through the above steps S601-S603, concrete private information is obtained by display collection and private knowledge is generated; further, the terminal device can expand the private knowledge set based on the collected data containing multiple private entities and the generated private knowledge by executing the above steps S604 and S605, so as to generate more high-confidence concrete private knowledge with as little display collection as possible and using a small amount of generated high-confidence private knowledge within the scope permitted by privacy security. In some examples, after executing the above steps, the terminal device can further determine whether to traverse each private entity in the minimum private knowledge set. If all private entities are traversed, that is, the features corresponding to all private entities are generated, then the construction of the minimum private knowledge set is completed. Otherwise, the above steps S604-S605 are repeated until all private entities are traversed and the constructed minimum private knowledge set is stored in the terminal device.

[0176] The following specifically describes the process of acquiring data corresponding to the first private entity in an implicit collection manner when constructing the private knowledge set.

[0177] FIG9 shows a flowchart of an implicit collection method according to an embodiment of the present application. The method can be executed on a terminal device, for example, on the mobile phone 101 in FIG1 above; as shown in FIG9, the method may include the following steps:

[0178] Step S901: Acquire initial data including a first private entity.

[0179] For example, the initial data may be data containing common sense information, such as wedding photos, family photos, selfies, etc. The terminal device may infer the data in the terminal device based on common sense to obtain initial data containing a first private entity. For example, if the user is a male user, wedding photos typically contain two private entities, "I" and "wife." The terminal device may use a model (such as a sub-model corresponding to the image modality) to extract features of images in the local gallery, thereby filtering out wedding photos from the local gallery. For another example, selfies typically contain the private entity "I." The terminal device may use a model to extract features of images in the local gallery, thereby filtering out selfies from the local gallery.

[0180] Exemplarily, the initial data obtained in this step includes concrete private information, so that the initial data including the first private entity can be used as data corresponding to the first private entity; thereby, the data corresponding to the first private entity is obtained without the user's awareness.

[0181] It is understandable that, since the user has not confirmed the data, the confidence level of the initial data obtained in this step is generally low. For example, the terminal device may contain multiple selfies of different people, which may contain data corresponding to the private entity "I" (i.e., the user's selfies) and may also contain data corresponding to other private entities (i.e., selfies of non-users). The terminal device can further filter and optimize the obtained initial data by executing the following steps S902-S903, thereby obtaining data corresponding to the first private entity with high confidence.

[0182] Step S902: Display at least one search result in response to a second search request of the user, wherein the second search request includes the first private entity, and the at least one search result corresponds to a different modality than the second search request.

[0183] Exemplarily, after receiving the user's second search request, the terminal device may process the second search request according to the model in the terminal device to obtain at least one search result, and may present the at least one search result to the user. As an example, the at least one search result may include the initial data containing the first private entity.

[0184] Step S903: Based on a selection operation of the user in at least one search result, data corresponding to the first private entity is obtained.

[0185] Exemplarily, the user can select the result he wants to search from at least one search result. Since the second search request includes the first private entity, the search result selected by the user includes the first private entity; thus, the terminal device can obtain the data corresponding to the first private entity, that is, highly confident concrete private information, based on the search result selected by the user.

[0186] Exemplarily, the terminal device may update the confidence level of the initial data based on a search result selected by the user. As an example, based on the user selecting the initial data from at least one search result, the terminal device may use the initial data as the data corresponding to the first private entity. For example, if the user selects the initial data, it indicates that the confidence level of the initial data is high, and the terminal device may use the initial data as the data corresponding to the first private entity. If the user does not select the initial data, it indicates that the confidence level of the initial data is still low, and the terminal device may not use the initial data as the data corresponding to the first private entity.

[0187] As another example, when the at least one search result does not include the above-mentioned initial data containing the first private entity, or when the at least one search result includes the initial data and the user does not select the initial data; the confidence level of the search result selected by the user is higher, then the terminal device can determine the search result selected by the user as the data corresponding to the first private entity.

[0188] For example, taking the mobile phone 101 in FIG. 1( a ) as the terminal device, and the first private entity being "I," as an example, FIG. 10( a )-( c ) illustrate a schematic diagram of an implicit data collection method according to an embodiment of the present application. Mobile phone 101 can filter out selfies from its local gallery. As shown in FIG. 10( a ), the second search request can be the keyword "my photos." The user can enter the keyword "my photos" into the search box of mobile phone 101's local gallery. As shown in FIG. 10( b ), mobile phone 101 processes the keyword using the sub-model corresponding to the image modality, searches for the user's photos in the gallery, and displays the searched user photos on the screen. As shown in FIG. 10( c ), the user can select a photo of their own from the photos displayed on the screen of mobile phone 101, for example, selecting Picture 2. If the filtered selfies include Picture 2, mobile phone 101 can use Picture 2 as the data corresponding to the private entity "I." In this way, the user's selection operation increases the confidence level of Picture 2, thereby obtaining highly confident, concrete private information.

[0189] In this way, through the above steps S901-S903, low-confidence concrete private information (i.e., initial data) is automatically obtained, and further in response to a second search request actively triggered by the user, at least one search result is displayed, and based on the user's operation of selecting the above-mentioned initial data in at least one search result, the initial data is used as the data corresponding to the first private entity, thereby improving the confidence of the concrete private information, and thereby obtaining the data corresponding to the first private entity in an implicit collection method that the user is unaware of.

[0190] As an example, through the above-mentioned implicit collection method, all or most of the concrete private information required to build a private knowledge set can be collected, which can serve as a backup for the above-mentioned explicit collection method. Even without explicit collection, the collection of concrete private information with high confidence can be completed.

[0191] Furthermore, the terminal device may also execute step S904, input the data corresponding to the first private entity into the model, and generate features corresponding to the first private entity; thereby generating private knowledge from the highly confident concrete private information obtained by the above implicit collection method.

[0192] Possible implementations of step S904 may refer to the relevant description in the above step S502.

[0193] Furthermore, the terminal device may also obtain features corresponding to other private entities in the private knowledge set by executing the following steps S905-S906, and automatically expand the private knowledge set, thereby completing the construction of the private knowledge set.

[0194] Step S905: Acquire data including the first private entity and the second private entity.

[0195] Step S906: Obtain features corresponding to the second private entity based on the data including the first private entity and the second private entity and the features corresponding to the first private entity.

[0196] The specific implementation process of steps S905-S906 can refer to the relevant descriptions in steps S503-S504 in Figure 5 above.

[0197] In this way, through the above steps S901-S904, concrete private information is acquired in an implicit collection manner and private knowledge is generated. Furthermore, the terminal device can expand the private knowledge set based on the collected data containing multiple private entities and the generated private knowledge by executing the above steps S905 and S906, thereby generating more high-confidence concrete private knowledge with as little collection as possible and utilizing a small amount of generated high-confidence private knowledge within the scope permitted by privacy security. In some examples, after executing the above steps, the terminal device can further determine whether to traverse each private entity in the minimum private knowledge set. If all private entities are traversed, that is, features corresponding to all private entities are generated, the construction of the minimum private knowledge set is completed. Otherwise, the above steps S905-S906 are repeated until all private entities are traversed, and the constructed minimum private knowledge set is stored in the terminal device. Exemplarily, a triple storage method can be used, for example, <1#, 2#, husband-wife relationship>, <1#, 3#, father-son relationship>, etc.

[0198] The following describes in detail the model training method provided by the embodiment of the present application from the training side. It is understandable that the application side and the training side can correspond to the same device, that is, the model training and the search using the trained model can be performed on the same device; the application side and the training side can also correspond to different devices, that is, the model can be trained on one device and the trained model can be configured on another device for concrete search. For example, the model can be pre-trained on the server, and then the model can be transplanted to the terminal device to realize the concrete search on the terminal device.

[0199] FIG11 shows a flow chart of a model training method according to an embodiment of the present application. Exemplarily, the model training method can be applied to a server, such as a cloud server. As shown in FIG11 , the method may include the following steps:

[0200] S1101: Obtain a multimodal sample set and a first model.

[0201] Among them, the multimodal sample set may include: multimodal samples corresponding to the first private entity, multimodal samples corresponding to the first search request, and multimodal samples corresponding to the second search request; the first private entity represents an entity that has an association relationship with the user, the first search request includes the first private entity, and the second search request does not include the first private entity.

[0202] Exemplarily, the acquired multimodal sample set can be divided into a training set, a validation set, and a test set; wherein, the number of samples contained in each of the training set, the validation set, and the test set can be set according to demand and is not limited to this.

[0203] Exemplarily, a multimodal sample may include a sample of a first modality and a sample of a second modality. For example, a multimodal sample may include a sample of a text modality and a sample of an image modality. As an example, a multimodal sample set may include a picture-text pair data set, wherein the picture-text pair data set includes multiple picture-text pair data, wherein the data of the text modality corresponds to the data of the image modality in each picture-text pair data. For example, "I" in the text modality (i.e., the text-I describing the user's photo) and "I" in the image modality (i.e., the user's photo) may be used as a picture-text pair data.

[0204] As an example, the multimodal sample corresponding to the first private entity may be the image-text pair data corresponding to the first private entity. For example, the first private entity may be "I", and the multimodal sample corresponding to the first private entity may be "I" in text mode and "I" in image mode. The multimodal sample corresponding to the first search request may be the image-text pair data corresponding to the first search request. For example, the multimodal sample corresponding to the first search request may be "I play badminton" in text mode (i.e., the keyword describing the photo of the user playing badminton - I play badminton) and "I play badminton" in image mode (i.e., the photo of the user playing badminton). The multimodal sample corresponding to the second search request may be the public image-text pair data; for example, the multimodal sample corresponding to the second search request may be "other people play badminton" in text mode (i.e., the keyword describing the photo of other people other than the user playing badminton - other people play badminton) and "other people play badminton" in image mode (i.e., the photo of other people other than the user playing badminton).

[0205] Exemplarily, the first model may be a trained multimodal model, for example, it may be a multimodal model trained using a public data set, which can be used to generate similar features for the same non-private entity in samples of different modalities; wherein, the training process of the first model can refer to the existing technology and will not be repeated here. Exemplarily, the first model may include: a sub-model corresponding to the first modality and a sub-model corresponding to the second modality. For example, the first modality may be an image modality, and the second modality may include a text modality; wherein, the sub-model corresponding to the text modality is used to encode the data of the text modality and generate corresponding features; the sub-model corresponding to the picture modality is used to encode the data of the image modality and generate corresponding features.

[0206] Exemplarily, a public image-text pair dataset can be used to divide the dataset into a training set, a validation set, and a test set to train the multimodal model. When the accuracy of the test set reaches a preset value, the same representation of the public image-text pairs in high-dimensional space, that is, the same high-dimensional space features, can be obtained. At this time, the training is stopped and the model is saved to obtain a trained multimodal model, that is, the first model.

[0207] S1102: Use the multimodal sample set to train the first model to obtain a second model.

[0208] It is understandable that, since the first model is obtained by training with a public data set, the first model has the ability to recognize general entities, for example, it can recognize the entity badminton; and for private entities such as "I", "wife" or "Zhang San", since the public data set does not distinguish private entities for different users, the first model trained with the public data set is only able to recognize these entities as "people" and cannot accurately correspond these entities to specific images in reality. In this step, a multimodal sample set containing private information (i.e., multimodal samples corresponding to the first private entity) is used to fine-tune the above-mentioned first model, thereby improving the expressiveness of the model. The obtained second model itself has the ability to carry concrete information, and is used to generate similar features for the same private entity in samples of different modalities, so that the private entity can be accurately matched to a specific image in reality. The second model can be used as the model in Figure 2 above for concrete search, greatly improving the search effect and efficiency.

[0209] Exemplarily, the multimodal samples corresponding to the first private entity, the multimodal samples corresponding to the first search request, and the multimodal samples corresponding to the second search request can be used to train the first model through comparative learning to obtain the second model; wherein, the multimodal samples corresponding to the first search request can be used as positive examples, and the multimodal samples corresponding to the second search request can be used as (difficult) negative examples.

[0210] In an embodiment of the present application, during the training of the first model, fine-tuning is performed based on a certain "private" data set, which enables the trained model (i.e., the second model) to generate similar features for the same private entity in samples of different modalities, and has the ability to infer the features of public information and visualize private information; and then the second model can be used for visualized search, greatly improving the search effect and efficiency.

[0211] The training process in step S1102 is described in detail below.

[0212] FIG12 is a flowchart of a model training method according to an embodiment of the present application. As shown in FIG12 , the above step S1102 may include the following steps:

[0213] S11021. Process the multimodal sample corresponding to the first private entity and the multimodal sample corresponding to the first search request through the first model to obtain a fusion feature.

[0214] For example, the server can obtain a multimodal sample set corresponding to a private entity, which can be used as a "private" dataset. For example, a data set of image-text pairs corresponding to the private entity can be obtained, where each image-text pair can include an image and a label of the private entity in the image.

[0215] Exemplarily, the server may parse the first search request, obtain the private entity contained therein (ie, the first private entity), and acquire the modal sample corresponding to the first private entity.

[0216] As an example, the server may input a sample of the first modality corresponding to the first private entity and a sample of the second modality corresponding to the first search request into the first model to obtain a fusion feature.

[0217] For example, the sample of the second modality corresponding to the first search request can be the keyword "I play badminton", and the private entity contained in the first search request can be "I", that is, the first private entity is "I", so the "I" of the image modality (that is, the sample of the first modality corresponding to the first private entity) and the keyword "I play badminton" can be input into the first model to obtain the fusion feature.

[0218] As an example, the sample of the second modality corresponding to the first search request may be the keyword "I play badminton". By parsing the keyword, the server can determine that the private entity contained in the first search request is "I", that is, the first private entity is "I", so that the "I" of the image modality can be input into the sub-model corresponding to the image modality. The sub-model corresponding to the image modality infers the features corresponding to the "I" of the image modality, that is, the features of the private entity "I" in the image.

[0219] In one possible implementation, the server may input the multimodal sample corresponding to the first search request into the first model to generate a first feature; input the multimodal sample corresponding to the first private entity into the first model to generate a second feature; and fuse the first and second features to obtain a fused feature. A detailed description of this implementation can be found in step S20301 of FIG. 3 above and will not be repeated here.

[0220] Exemplarily, the server may input the sample of the second modality corresponding to the first search request into the sub-model corresponding to the second modality to obtain the first feature; input the sample of the first modality corresponding to the first private entity into the sub-model corresponding to the first modality to obtain the second feature; and then fuse the first feature and the second feature to obtain a fused feature. For example, if the sample of the text modality corresponding to the first search request is the keyword "I play badminton", then the keyword "I play badminton" can be input into the sub-model corresponding to the text modality, and the sub-model corresponding to the text modality infers the feature corresponding to the keyword "I play badminton" (i.e., the first feature); if the sample of the image modality corresponding to the first search request is a photo of "me", then the photo of "me" can be input into the sub-model corresponding to the image modality, and the sub-model corresponding to the image modality infers the feature corresponding to the photo of "me" (i.e., the second feature), and then the two features are fused to obtain a fused feature.

[0221] In this way, during the training process of the first model, the first model is used to obtain fused features; as an example, private features (i.e., the second features) and public features (i.e., the first features) are generated, and the generated public features and private features are fused. The fused features can simultaneously represent public information and concrete private information, so that the first model can take into account the learning of public information and private information.

[0222] S11022. Train the first model based on the fused features, the features corresponding to the first search request, and the features corresponding to the second search request to obtain a second model.

[0223] Among them, the features corresponding to the first search request are obtained by the first model processing the multimodal samples corresponding to the first search request, and the features corresponding to the second search request are obtained by the first model processing the multimodal samples corresponding to the second search request.

[0224] Exemplarily, the features corresponding to the first search request are obtained by processing the samples of the first modality corresponding to the first search request by the first model, and the features corresponding to the second search request are obtained by processing the samples of the first modality corresponding to the second search request by the first model. As an example, the samples of the first modality corresponding to the first search request may be photos of "I play badminton", and the samples of the first modality corresponding to the second search request may be photos of "other people playing badminton". The server may input the photos of "I play badminton" into the sub-model corresponding to the image modality to obtain the features corresponding to "I play badminton" in the image modality, and input the photos of "other people playing badminton" into the sub-model corresponding to the image modality to obtain the features corresponding to "other people playing badminton" in the image modality.

[0225] For example, the first model can be trained through comparative learning using the fused features, the features corresponding to the first search request, and the features corresponding to the second search request to obtain a second model. The features corresponding to the first search request serve as the features of the positive example, and the features corresponding to the second search request serve as the features of the negative example. During the comparative learning process, the distance between the features of the positive example and the fused features gradually decreases, while the distance between the features of the negative example and the fused features gradually increases, until the features of the positive example and the fused features are aligned. That is, the distance between the features of the positive example and the fused features is close, such as the Euclidean distance or cosine distance, while the distance between the features of the positive example and the fused features is greater. It can be understood that the greater the Euclidean distance between two features, the greater the difference between the two features, i.e., the lower the similarity. When the test set accuracy reaches a preset value, training can be stopped and the model saved, resulting in a trained public-private fusion multimodal model, i.e., the second model. This second model can be used as the model in Figure 2 above. During the search process, the second model combines private knowledge to infer fused features that can represent public and private information, thereby achieving concrete search and greatly improving search effectiveness and efficiency.

[0226] Exemplarily, in the process of comparative learning, the loss function value can be determined based on the fused features, the features corresponding to the first search request, and the features corresponding to the second search request. For example, the loss function value of the positive example can be calculated by the difference between the fused features and the features corresponding to the first search request, and the loss function value of the negative example can be calculated by the difference between the fused features and the features corresponding to the second search request. The loss function value is obtained based on the loss function value of the positive example and the loss function value of the negative example. After obtaining the loss function value, the loss function value can be backpropagated, and the gradient descent algorithm is used to update the parameter values ​​in the first model. Through continuous training, the degree of similarity (such as Euclidean distance) between the fused features obtained by the first model and the features of the positive example continues to approach, while the degree of similarity with the features of the negative example gradually increases.

[0227] For example, FIG13 illustrates a flow chart of a model training method according to an embodiment of the present application. As shown in FIG13 , the first private entity may be "I," and the multimodal sample corresponding to the first private entity is a photo of "I" and the text label "I" corresponding to the photo; the multimodal sample corresponding to the first search request is a photo of "I" playing badminton and the text label "I play badminton" corresponding to the photo; the multimodal sample corresponding to the second search request is a photo of other people playing badminton and the text label "other people play badminton" corresponding to the photo; wherein the photo of "I" playing badminton serves as a positive example, and the photo of other people playing badminton serves as a negative example; during fine-tuning of the first model, the first model (e.g., the sub-model corresponding to the text modality) may process the text label "I play badminton" to obtain a first feature; the first model (e.g., the sub-model corresponding to the image modality) may process the photo of "I" to obtain feature A, i.e., the second feature; the photo of "I" playing badminton may be processed to obtain feature B; and the photo of other people playing badminton may be processed to obtain feature C. The first model fuses the first and second features to obtain a fused feature. Then, the first model is fine-tuned by combining feature B, feature C, and the fused feature. Through continuous fine-tuning, feature B gradually approaches the fused feature, while feature C gradually moves away from the fused feature, until the feature alignment between feature B and the fused feature is achieved, that is, the "I" in the text modality is aligned with the "I" in the image modality. When the accuracy of the test set reaches the preset value, the training can be stopped, the model can be saved, and the second model can be obtained. In this way, the model training process can take into account both the public information of the photos and the concrete private information. When searching in the gallery using the trained second model, the public and private information of the pictures in the gallery can be extracted, and the fused features can be inferred and generated. Then, combined with the private knowledge, pictures containing private information such as "I" and "wife" can be accurately searched.

[0228] In an embodiment of the present application, a multimodal model and private data are used to fuse private features with public features, and then align them with the corresponding multimodal features to train a public-private fusion multimodal model. In this way, the first model is fine-tuned using joint training of public and private features, which enables the trained model (i.e., the second model) to have the ability to infer concrete private information, thus solving the problem that a multimodal model trained only with a public data set cannot carry concrete private information.

[0229] Based on the same inventive concept as the above method embodiments, embodiments of the present application further provide a search device that can be used to implement the technical solutions described in the above method embodiments. For example, the steps of the method shown in Figures 2, 3, 5, 6, or 9 can be executed.

[0230] Figure 14 shows a structural diagram of a search device according to an embodiment of the present application. As shown in Figure 14, a receiving module 1401 is used to receive a first search request from a user, wherein the first search request includes a private entity; wherein the private entity represents an entity associated with the user; a determining module 1402 is used to determine, from among the features corresponding to at least one private entity, a feature corresponding to the private entity in the first search request; wherein the feature corresponding to the at least one private entity indicates data corresponding to the at least one private entity, and the data corresponding to the at least one private entity corresponds to a different modality than the first search request; a searching module 1403 is used to obtain search results based on the first search request and the features corresponding to the private entity in the first search request; wherein the search results correspond to a different modality than the first search request; and a display module 1404 is used to display the search results.

[0231] In an embodiment of the present application, among the features corresponding to at least one private entity, the features corresponding to the private entity in the user's search request are determined, and the search results are obtained based on the features corresponding to the private entity and the first search request. Since the features corresponding to at least one private entity indicate the data corresponding to at least one private entity, the data corresponding to at least one private entity can correspond to the same modality as the search results, so that the private entity in the search request can be corresponded to the specific image in reality (that is, the specific image of the private entity in the search results), thereby realizing a figurative search; solving the problem that the terminal device cannot perform a figurative search during the search process, improving the search effect and search efficiency, and enhancing the user's search experience. As an example, the model can be used to infer the private entity in the first search request and the first search request to obtain public-private fusion features, and then the search results can be determined in the database through the public-private fusion features. In addition, there is no need for a closed tag set, and it is highly flexible; it can meet the user's diversified search requests, and it is more scalable; for different terminal devices, there is no need to replace the model, and maintenance is easier.

[0232] In one possible implementation, the search module 1403 is further used to: process the first search request and the features corresponding to the private entities in the first search request through a model to generate a fusion feature; the fusion feature indicates the first search request; and use the data corresponding to the features in the database that match the fusion feature as the search result.

[0233] In one possible implementation, the device further includes: a generation module for obtaining data corresponding to a first private entity; the first private entity is any one of the at least one private entity; and processing the data corresponding to the first private entity to generate features corresponding to the first private entity.

[0234] In a possible implementation, the generation module is further configured to: send a first prompt message to the user, where the first prompt message is used to prompt the user to select data containing the first private entity; and obtain data corresponding to the first private entity in response to the user's selection operation.

[0235] In one possible implementation, the generation module is further used to: obtain initial data containing the first private entity; display at least one search result in response to a second search request from a user; the second search request contains the first private entity, and the at least one search result corresponds to a different modality from the second search request; the at least one search result includes the initial data; and based on the user's operation of selecting the initial data in the at least one search result, use the initial data as data corresponding to the first private entity.

[0236] In one possible implementation, the generation module is further used to: obtain data containing the first private entity and the second private entity; wherein the second private entity is any private entity among the at least one private entity except the first private entity; and obtain features corresponding to the second private entity based on the data containing the first private entity and the second private entity, and features corresponding to the first private entity.

[0237] In a possible implementation, the private entity includes: a title, a name, or a nickname associated with the user.

[0238] The technical effects and specific descriptions of the search device shown in FIG. 14 and its various possible implementation methods can be found in the above-mentioned search method, which will not be repeated here.

[0239] Based on the same inventive concept of the above-mentioned method embodiment, an embodiment of the present application also provides another search device, which includes: a first display module, used to display a first display interface; the first display interface includes a search portal; a second display module, used to display a second display interface in response to the user inputting a keyword at the search portal, the second display interface including search results corresponding to the keyword; wherein the keyword includes a private entity; the search results are obtained based on the features corresponding to the keyword and the private entity; the search results and the keyword correspond to different modalities.

[0240] In an embodiment of the present application, in response to the user entering a keyword in the search entrance of the first display interface, a second display interface is displayed; since the search results are obtained based on the keyword and the characteristics corresponding to the private entity in the search request, the private entity in the keyword can be corresponded to the specific image in reality (that is, the specific image of the private entity in the search results), thereby realizing a concrete search; this solves the problem that the terminal device cannot perform a concrete search during the search process, improves the search effect and search efficiency, and enhances the user's search experience.

[0241] In one possible implementation, the device also includes: a third display module, used to display a third display interface, the third display interface including a first prompt information and at least one data; the first prompt information is used to prompt the user to select data containing the first private entity; a fourth display module, used to display a fourth display interface in response to the user's selection operation in the at least one data, the fourth display interface including first data and a first identifier, the first identifier being used to indicate that the first data is the data selected by the user.

[0242] In one possible implementation, the second display interface also includes: a second prompt message, the second prompt message is used to prompt the user to confirm whether the first search result is the search result expected by the user; the device also includes: a fifth display module, used to display a fifth display interface in response to the user's confirmation operation, the fifth display interface includes the first search result and a second identifier, the second identifier is used to indicate that the first search result is the search result confirmed by the user.

[0243] Based on the same inventive concept of the above method embodiments, embodiments of the present application further provide a model training device, which can be used to implement the technical solutions described in the above method embodiments. For example, the steps of the method shown in FIG. 11 or FIG. 12 can be executed.

[0244] Figure 15 shows a structural diagram of a model training device according to an embodiment of the present application. As shown in Figure 15, the device includes: an acquisition module 1501, which is used to acquire a multimodal sample set and a first model; wherein the multimodal sample set includes: multimodal samples corresponding to a first private entity, multimodal samples corresponding to a first search request, and multimodal samples corresponding to a second search request; the first private entity represents an entity associated with a user, the first search request includes the first private entity, and the second search request does not include the first private entity; and a training module 1502, which is used to train the first model using the multimodal sample set to obtain a second model.

[0245] In an embodiment of the present application, during the training of the first model, fine-tuning is performed based on a certain "private" data set, which enables the trained model (i.e., the second model) to generate similar features for the same private entity in samples of different modalities, and has the ability to infer the features of public information and visualize private information; and then the second model can be used for visualized search, greatly improving the search effect and efficiency.

[0246] In one possible implementation, the training module 1502 is further used to: process the multimodal samples corresponding to the first private entity and the multimodal samples corresponding to the first search request through the first model to obtain fused features; train the first model according to the fused features, the features corresponding to the first search request, and the features corresponding to the second search request to obtain the second model; wherein the features corresponding to the first search request are obtained by processing the multimodal samples corresponding to the first search request by the first model, and the features corresponding to the second search request are obtained by processing the multimodal samples corresponding to the second search request by the first model.

[0247] In one possible implementation, the training module 1502 is further used to: input the multimodal sample corresponding to the first search request into the first model to generate a first feature; input the multimodal sample corresponding to the first private entity into the first model to generate a second feature; and fuse the first feature and the second feature to obtain the fused feature.

[0248] In a possible implementation, the multimodal samples include samples of the first modality and samples of the second modality; the features corresponding to the first search request are obtained by processing the samples of the first modality corresponding to the first search request by the first model, and the features corresponding to the second search request are obtained by processing the samples of the first modality corresponding to the second search request by the first model; the training module 1502 is further used to: input the samples of the first modality corresponding to the first private entity and the samples of the second modality corresponding to the first search request into the first model to obtain the fused features

[0249] The technical effects and specific descriptions of the model training device shown in Figure 15 and its various possible implementation methods can be found in the above-mentioned model training method, which will not be repeated here.

[0250] It should be understood that the division of the modules in the above search device and model training device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. In addition, the modules in the device can be implemented in the form of a processor calling software; for example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or implement the functions of the modules of the device, wherein the processor is, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory is a memory inside the device or a memory outside the device. Alternatively, the modules in the device can be implemented in the form of hardware circuits, and the functions of some or all modules can be realized by designing the hardware circuits. The hardware circuit can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC), which realizes the functions of some or all of the above modules by designing the logical relationship of the components in the circuit. For another example, in another implementation, the hardware circuit can be implemented by a programmable logic device (PLD). Taking a field programmable gate array (FPGA) as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by configuring the configuration file, thereby realizing the functions of some or all of the above modules. All modules of the above devices can be implemented in the form of software called by the processor, or in the form of hardware circuits, or in part by the form of software called by the processor, and the rest by hardware circuits.

[0251] In an embodiment of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor may be a circuit with the ability to read and execute instructions, such as a CPU, a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a neural-network processing unit (NPU), a tensor processing unit (TPU), etc. In another implementation, the processor may implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or reconfigurable, such as a hardware circuit implemented by an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above modules.

[0252] It can be seen that each module in the above apparatus can be one or more processors (or processing circuits) configured to implement the above embodiment methods, such as: CPU, GPU, NPU, TPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms. In addition, each module in the above apparatus can be fully or partially integrated together, or can be implemented independently, without limitation.

[0253] The present application also provides an electronic device comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the method of the above embodiment when executing the instructions. For example, the method may include the steps of the method shown in Figures 2, 3, 5, 6, or 9, or the steps of the method shown in Figures 11 or 11.

[0254] As an example, FIG16 shows a schematic structural diagram of an electronic device 100 according to an embodiment of the present application. The electronic device 100 may include at least one terminal device selected from a mobile phone, a foldable electronic device, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) device, a wearable device, an in-vehicle device, a smart home device, or a smart city device. The embodiment of the present application does not impose any special restrictions on the specific type of the electronic device 100.

[0255] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) connector 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0256] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0257] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0258] The processor can generate an operation control signal based on the instruction opcode and the timing signal to complete the control of instruction fetching and execution. For example, the processor can perform the steps of the method shown in Figures 2, 3, 5, 6 or 9 above.

[0259] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 may be a cache memory. This memory can store instructions or data that have been used or are frequently used by processor 110. If processor 110 needs to use the instruction or data, it can directly access it from this memory. This avoids repeated accesses, reduces processor 110 latency, and thus improves system efficiency. As an example, this memory may store models or private knowledge sets.

[0260] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface. The processor 110 may be connected to modules such as a touch sensor, an audio module, a wireless communication module, a display, and a camera through at least one of the above interfaces.

[0261] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0262] The USB connector 130 is an interface that complies with USB standards and can be used to connect the electronic device 100 and peripheral devices. Specifically, it can be a Mini USB connector, a Micro USB connector, a USB Type-C connector, etc. The USB connector 130 can be used to connect to a charger to charge the electronic device 100, or to connect to other electronic devices to transfer data between the electronic device 100 and other electronic devices. It can also be used to connect headphones to output audio stored in the electronic device through the headphones. This connector can also be used to connect to other electronic devices, such as VR devices.

[0263] The charging management module 140 is configured to receive charging input from a charger. The charger may be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 may receive charging input from the wired charger via the USB interface 130.

[0264] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 and provides power to the processor 110, the internal memory 121, the display 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance).

[0265] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.

[0266] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.

[0267] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.

[0268] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.

[0269] The wireless communication module 160 can provide wireless communication solutions applied to the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), Bluetooth low energy (BLE), ultra wide band (UWB), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0270] In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with a network and other electronic devices through wireless communication technology. The wireless communication technology may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).

[0271] Electronic device 100 can implement display functions using a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0272] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or more display screens 194. Exemplarily, the display screen 194 can be used to display a display interface, search requests, search results, or prompt information, etc.

[0273] The electronic device 100 can realize the camera function through the camera module 193, ISP, video codec, GPU, display screen 194, application processor AP, neural network processor NPU, etc.

[0274] The camera module 193 can be used to collect color image data and depth data of the subject. The ISP can be used to process the color image data collected by the camera module 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then transmitted to the ISP for processing and converted into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise, brightness, and skin color. The ISP can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be provided in the camera module 193.

[0275] In some embodiments, the camera module 193 may be composed of a color camera module and a 3D sensing module. In some embodiments, the 3D sensing module may be a (time of flight, TOF) 3D sensing module or a structured light 3D sensing module. In other embodiments, the camera module 193 may also be composed of two or more cameras. In some embodiments, the electronic device 100 may include one or more camera modules 193. Specifically, the electronic device 100 may include one front camera module 193 and one rear camera module 193. Among them, the front camera module 193 can generally be used to collect the color image data and depth data of the photographer himself facing the display screen 194, and the rear camera module can be used to collect the color image data and depth data of the subject (such as people, scenery, etc.) facing the photographer.

[0276] In some embodiments, the CPU or GPU or NPU in the processor 110 can be based on Huawei's PoissonEngine_VS vector engine, or open source vector engines such as Faiss, Annoy, Milvus, Vearch, etc.; perform private knowledge reasoning, build a private knowledge set, and store the private knowledge set in the data; and can also infer keywords based on user search requests to obtain fusion features corresponding to the keywords.

[0277] Digital signal processors are used to process digital signals and can also process other digital signals.

[0278] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. This allows electronic device 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.

[0279] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU can enable intelligent cognitive applications in electronic device 100, such as image recognition, face recognition, speech recognition, and text comprehension.

[0280] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be saved on the external memory card or transferred from the electronic device to the external memory card.

[0281] The internal memory 121 can be used to store computer executable program code, which includes instructions. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. As an example, a private knowledge set and a database may be stored. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional methods or data processing of the electronic device 100 by running instructions stored in the internal memory 121, and / or instructions stored in a memory provided in the processor. Exemplarily, the methods shown in Figures 2, 3, 5, 6 or 9 above may be executed.

[0282] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.

[0283] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.

[0284] The speaker 170A, also called a "speaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or output audio signals for hands-free calls through the speaker 170A.

[0285] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a call or a voice message, the user can place the receiver 170B close to the ear to hear the voice.

[0286] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the source of sound, realize directional recording function, etc.

[0287] The headphone jack 170D is used to connect a wired headphone and can be the USB interface 130 or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0288] The pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, the pressure sensor 180A can be set on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. A capacitive pressure sensor can be a device comprising at least two parallel plates with conductive material. When force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure based on the change in capacitance. When a touch operation is applied to the display screen 194, the electronic device 100 detects the intensity of the touch operation based on the pressure sensor 180A. The electronic device 100 can also calculate the position of the touch based on the detection signal of the pressure sensor 180A. In some embodiments, touch operations acting on the same touch position but with different touch operation intensities can correspond to different operation instructions.

[0289] The gyro sensor 180B can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., x, y, and z axes) can be determined by the gyro sensor 180B. The gyro sensor 180B can also be used for image stabilization.

[0290] The air pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates the altitude based on the air pressure value measured by the air pressure sensor 180C to assist in positioning and navigation.

[0291] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip cover. When the electronic device is a foldable electronic device, the magnetic sensor 180D can be used to detect whether the electronic device is folded or unfolded, or the folding angle.

[0292] The acceleration sensor 180E can detect the magnitude of acceleration in various directions (generally three axes) of the electronic device 100. When the electronic device 100 is stationary, the magnitude and direction of gravity can be detected.

[0293] The distance sensor 180F is used to measure distance. The electronic device 100 can measure distance using infrared or laser. In some embodiments, when shooting a scene, the electronic device 100 can use the distance sensor 180F to measure distance to achieve fast focusing.

[0294] The proximity light sensor 180G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared light emitting diode. The electronic device 100 emits infrared light outward through the LED.

[0295] Ambient light sensor 180L can be used to sense ambient light brightness. Electronic device 100 can adaptively adjust the brightness of display screen 194 based on the perceived ambient light brightness. Ambient light sensor 180L can also be used to automatically adjust white balance when taking photos. Ambient light sensor 180L can also work with proximity light sensor 180G to detect whether electronic device 100 is obscured, such as when the device is in a pocket.

[0296] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to implement fingerprint unlocking, access application locks, fingerprint photography, fingerprint call answering, etc.

[0297] Temperature sensor 180J is used to detect temperature. In some embodiments, electronic device 100 utilizes the temperature detected by temperature sensor 180J to implement a temperature management strategy. For example, when the temperature detected by temperature sensor 180J exceeds a threshold, electronic device 100 may reduce processor performance to reduce power consumption and implement thermal protection.

[0298] The touch sensor 180K is also called a "touch-sensitive device." The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also called a "touch screen." The touch sensor 180K is used to detect touch operations applied thereto or in the vicinity thereof. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operations can be provided via the display screen 194. In other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, at a location different from that of the display screen 194.

[0299] The bone conduction sensor 180M can obtain a vibration signal. In some embodiments, the bone conduction sensor 180M can obtain a vibration signal of a vibrating bone in a human vocal part.

[0300] The buttons 190 may include a power button, a volume button, etc. The buttons 190 may be mechanical buttons or touch buttons. The electronic device 100 may receive key inputs and generate key signal inputs related to user settings and function control of the electronic device 100.

[0301] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 194, motor 191 can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0302] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, etc.

[0303] The SIM card interface 195 is used to connect a SIM card. The SIM card can be connected to or disconnected from the electronic device 100 by inserting it into or removing it from the SIM card interface 195. The electronic device 100 can support one or more SIM card interfaces. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to implement functions such as calls and data communications.

[0304] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. In the embodiment of the present application, the Android system with a layered architecture is used as an example to illustrate the software structure of the electronic device 100.

[0305] FIG17 shows a software structure block diagram of the electronic device 100 according to an embodiment of the present application.

[0306] A layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0307] The application layer can include a series of application packages.

[0308] As shown in FIG17 , the application package may include applications such as phone, camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and short message.

[0309] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0310] As shown in FIG17 , the application framework layer may include a window manager, a content provider, a view system, a telephony manager, a resource manager, a notification manager, and the like.

[0311] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.

[0312] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, private knowledge sets, etc.

[0313] The view system includes visual controls, such as controls for displaying text, controls for displaying images, and controls for displaying a gallery user interface (UI). The view system can be used to build applications. A display interface can be composed of one or more views. For example, a display interface that includes a text notification icon can include a view that displays text and a view that displays images. As an example, it can be used to provide a display interface for image viewing, search entry, and search results.

[0314] The phone manager is used to provide communication functions for terminal devices, such as call status management (including answering, hanging up, etc.).

[0315] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.

[0316] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically, without requiring user interaction. For example, the Notification Manager can be used to notify users of completed downloads and message reminders. The Notification Manager can also display notifications in the top status bar of the system as icons or scrolling text, such as notifications from background applications, or as dialog windows on the screen. Examples include displaying text messages in the status bar, emitting alert sounds, vibrating the device, or flashing indicator lights.

[0317] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for scheduling and management of the Android system.

[0318] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.

[0319] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.

[0320] The system library can include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.

[0321] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.

[0322] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0323] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0324] A 2D graphics engine is a drawing engine for 2D drawings.

[0325] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.

[0326] As another example, Figure 18 shows a structural schematic diagram of another electronic device according to an embodiment of the present application. Exemplarily, the electronic device may be a server; as shown in Figure 18, the electronic device may include: at least one processor 1601, a communication line 1602, a memory 1603 and at least one communication interface 1604.

[0327] Processor 1601 can be a general-purpose central processing unit, a microprocessor, a specific application integrated circuit, or one or more integrated circuits used to control the execution of the program of the present application; processor 1601 can also include a heterogeneous computing architecture of multiple general-purpose processors, for example, it can be a combination of at least two of CPU, GPU, microprocessor, DSP, ASIC, and FPGA; as an example, processor 1601 can be CPU+GPU or CPU+ASIC or CPU+FPGA.

[0328] Communication link 1602 may include a pathway for transmitting information between the aforementioned components.

[0329] The communication interface 1604 uses any transceiver or other device for communicating with other devices or communication networks, such as Ethernet, RAN, wireless local area networks (WLAN), etc.

[0330] Memory 1603 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program codes in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory can be independent and connected to the processor via a communication line 1602. The memory can also be integrated with the processor. The memory provided in the embodiment of the present application can generally have non-volatility. Among them, the memory 1603 is used to store computer-executable instructions for executing the solution of the present application, and is controlled by the processor 1601 for execution. The processor 1601 is used to execute the computer-executable instructions stored in the memory 1603, thereby implementing the method provided in the above embodiments of the present application; illustratively, the steps of the method shown in Figure 11 or Figure 12 can be implemented.

[0331] Optionally, the computer-executable instructions in the embodiments of the present application may also be referred to as application code, which is not specifically limited in the embodiments of the present application.

[0332] Exemplarily, processor 1601 may include one or more CPUs, such as CPU0 in FIG18 ; processor 1601 may also include a CPU and any one of a GPU, an ASIC, and an FPGA, such as CPU0+GPU0, CPU0+ASIC0, or CPU0+FPGA0 in FIG18 . In some examples, an NPU may also be included. The GPU and NPU may provide the computing power required for machine learning and neural network operators.

[0333] For example, an electronic device may include multiple processors, such as processor 1601 and processor 1607 in FIG18 . Each of these processors may be a single-core (single-CPU) processor, a multi-core (multi-CPU) processor, or a heterogeneous computing architecture including multiple general-purpose processors. A processor herein may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0334] In a specific implementation, as an embodiment, the electronic device may further include an output device 1605 and an input device 1606. The output device 1605 communicates with the processor 1601 and can display information in a variety of ways. For example, the output device 1605 can be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. For example, it can be a display device such as a vehicle-mounted HUD, an AR-HUD, a display, etc. The input device 1606 communicates with the processor 1601 and can receive user input in a variety of ways. For example, the input device 1606 can be a mouse, a keyboard, a touch screen device, or a sensing device, etc.

[0335] Embodiments of the present application provide a computer-readable storage medium having computer program instructions stored thereon. When executed by a processor, the computer program instructions implement the methods of the above-described embodiments. For example, the computer program instructions may implement the steps of the methods shown in Figures 2, 3, 5, 6, or 9, or the steps of the methods shown in Figures 11 or 12.

[0336] Embodiments of the present application provide a computer program product, which may include, for example, computer-readable code or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer program product is executed on a computer, the computer executes the method described in the above embodiments. For example, the computer program product may execute the steps of the method shown in Figures 2, 3, 5, 6, or 9, or the steps of the method shown in Figures 11 or 12.

[0337] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0338] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0339] The computer program instructions for performing the operation of the present application can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data or source code or object code written in any combination of one or more programming languages, wherein the programming language includes object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or executed completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer by any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (such as by using an Internet service provider to connect to the Internet). In certain embodiments, by utilizing the state information of computer-readable program instructions to personalize electronic circuits, such as programmable logic circuits, field programmable gate arrays (FPGAs) or programmable logic arrays (PLAs), the electronic circuits can execute computer-readable program instructions, thereby realizing various aspects of the present application.

[0340] Various aspects of the present application are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0341] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0342] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0343] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the system, method and computer program product according to multiple embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a special hardware-based system that performs the function or action of the specification, or can be implemented by a combination of special hardware and computer instructions.

[0344] While various embodiments of the present application have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A search method, characterized in that: The method comprises: Receive a first search request from a user, wherein the first search request includes a private entity; wherein the private entity represents an entity associated with the user; Determining, from among the features corresponding to the at least one private entity, a feature corresponding to the private entity in the first search request; wherein the feature corresponding to the at least one private entity indicates data corresponding to the at least one private entity, and the data corresponding to the at least one private entity corresponds to a different modality than the first search request; Obtaining search results based on the first search request and features corresponding to the private entity in the first search request; the search results and the first search request correspond to different modalities; The search results are displayed.

2. The method according to claim 1, characterized in that Obtaining search results according to the first search request and the features corresponding to the private entities in the first search request includes: Processing the first search request and features corresponding to the private entities in the first search request through a model to generate a fused feature; the fused feature indicates the first search request; The data corresponding to the feature matching the fusion feature in the database is used as the search result.

3. The method according to claim 1 or 2, characterized in that The method further comprises: Acquire data corresponding to a first private entity; the first private entity is any private entity among the at least one private entity; The data corresponding to the first private entity is processed to generate features corresponding to the first private entity.

4. The method according to claim 3, characterized in that The acquiring of data corresponding to the first private entity includes: Sending a first prompt message to the user, where the first prompt message is used to prompt the user to select data containing the first private entity; In response to a selection operation by the user, data corresponding to the first private entity is obtained.

5. The method according to claim 3, characterized in that The acquiring of data corresponding to the first private entity includes: Acquire initial data including the first private entity; In response to a second search request from the user, displaying at least one search result; the second search request includes the first private entity, the at least one search result corresponds to a different modality than the second search request; the at least one search result includes the initial data; Based on an operation of the user selecting the initial data from the at least one search result, the initial data is used as data corresponding to the first private entity.

6. The method according to any one of claims 3 to 5, characterized in that The method further comprises: Acquire data including the first private entity and a second private entity; wherein the second private entity is any private entity among the at least one private entity except the first private entity; The feature corresponding to the second private entity is obtained according to the data including the first private entity and the second private entity and the feature corresponding to the first private entity.

7. The method according to any one of claims 1 to 6, characterized in that The private entity includes: a title, a name, or a nickname associated with the user.

8. A search method, characterized in that: The method comprises: Displaying a first display interface; the first display interface includes a search entry; In response to the user inputting a keyword in the search portal, a second display interface is displayed, and the second display interface includes a first search result corresponding to the keyword; wherein the keyword includes a private entity; the first search result is obtained based on the characteristics corresponding to the keyword and the private entity; the first search result and the keyword correspond to different modalities.

9. The method according to claim 8, characterized in that The method further comprises: Displaying a third display interface, wherein the third display interface includes first prompt information and at least one data; the first prompt information is used to prompt the user to select data containing the first private entity; In response to a selection operation by the user in the at least one data, a fourth display interface is displayed, wherein the fourth display interface includes the first data and a first identifier, wherein the first identifier is used to indicate that the first data is the data selected by the user.

10. The method according to claim 8 or 9, characterized in that The second display interface further includes: second prompt information, the second prompt information being used to prompt the user to confirm whether the first search result is the search result expected by the user; The method further includes: displaying a fifth display interface in response to a confirmation operation of the user, wherein the fifth display interface includes the first search result and a second identifier, wherein the second identifier is used to indicate that the first search result is the search result confirmed by the user.

11. A model training method, characterized in that: The method comprises: Obtaining a multimodal sample set and a first model; wherein the multimodal sample set includes: multimodal samples corresponding to a first private entity, multimodal samples corresponding to a first search request, and multimodal samples corresponding to a second search request; the first private entity represents an entity associated with the user, the first search request includes the first private entity, and the second search request does not include the first private entity; The first model is trained using the multimodal sample set to obtain a second model.

12. The method according to claim 11, characterized in that The step of training the first model using the multimodal sample set to obtain a second model includes: Processing the multimodal sample corresponding to the first private entity and the multimodal sample corresponding to the first search request using the first model to obtain a fusion feature; The first model is trained based on the fused features, the features corresponding to the first search request, and the features corresponding to the second search request to obtain the second model; wherein, the features corresponding to the first search request are obtained by processing the multimodal samples corresponding to the first search request by the first model, and the features corresponding to the second search request are obtained by processing the multimodal samples corresponding to the second search request by the first model.

13. The method according to claim 12, characterized in that The processing of the multimodal sample corresponding to the first private entity and the multimodal sample corresponding to the first search request by the first model to obtain a fusion feature includes: Inputting the multimodal sample corresponding to the first search request into the first model to generate a first feature; inputting the multimodal sample corresponding to the first private entity into the first model to generate a second feature; The first feature and the second feature are fused to obtain the fused feature.

14. The method according to claim 12 or 13, characterized in that The multimodal samples include samples of a first modality and samples of a second modality; the features corresponding to the first search request are obtained by processing the samples of the first modality corresponding to the first search request by the first model, and the features corresponding to the second search request are obtained by processing the samples of the first modality corresponding to the second search request by the first model; The processing of the multimodal sample corresponding to the first private entity and the multimodal sample corresponding to the first search request by the first model to obtain a fusion feature includes: The sample of the first modality corresponding to the first private entity and the sample of the second modality corresponding to the first search request are input into the first model to obtain the fusion feature.

15. A search device, characterized in that: The device comprises: A receiving module, configured to receive a first search request from a user, wherein the first search request includes a private entity; wherein the private entity represents an entity associated with the user; a determining module, configured to determine, from among the features corresponding to the at least one private entity, a feature corresponding to the private entity in the first search request; wherein the feature corresponding to the at least one private entity indicates data corresponding to the at least one private entity, and the data corresponding to the at least one private entity corresponds to a different modality than the first search request; A search module, configured to obtain search results based on the first search request and features corresponding to the private entities in the first search request; the search results and the first search request correspond to different modalities; The display module is configured to display the search results.

16. The device according to claim 15, characterized in that The search module is further used to: process the first search request and the features corresponding to the private entities in the first search request through a model to generate a fusion feature; the fusion feature indicates the first search request; and use the data corresponding to the features in the database that match the fusion feature as the search result.

17. The device according to claim 15 or 16, characterized in that The device also includes: a generation module, configured to obtain data corresponding to a first private entity; the first private entity being any one of the at least one private entity; and processing the data corresponding to the first private entity to generate features corresponding to the first private entity.

18. The device according to claim 17, characterized in that The generating module is further configured to: send a first prompt message to the user, where the first prompt message is used to prompt the user to select data containing the first private entity; and obtain data corresponding to the first private entity in response to the user's selection operation.

19. The device according to claim 17, characterized in that The generating module is further configured to: obtain initial data including the first private entity; and display at least one search result in response to a second search request from a user, wherein the second search request includes the first private entity, and the at least one search result corresponds to a different modality than the second search request; The at least one search result includes the initial data; Based on an operation of the user selecting the initial data from the at least one search result, the initial data is used as data corresponding to the first private entity.

20. The device according to any one of claims 17 to 19, characterized in that The generation module is further configured to: obtain data including the first private entity and the second private entity; wherein the second private entity is any private entity among the at least one private entity other than the first private entity; and obtain features corresponding to the second private entity based on the data including the first private entity and the second private entity and features corresponding to the first private entity.

21. The device according to any one of claims 15 to 20, characterized in that The private entity includes: a title, a name, or a nickname associated with the user.

22. A search device, characterized in that: The device comprises: A first display module, configured to display a first display interface; the first display interface includes a search entry; The second display module is used to display a second display interface in response to the user inputting a keyword in the search portal, wherein the second display interface includes search results corresponding to the keyword; wherein the keyword includes a private entity; the search results are obtained based on the features corresponding to the keyword and the private entity; the search results and the keyword correspond to different modalities.

23. The device according to claim 22, characterized in that The device also includes: a third display module, used to display a third display interface, the third display interface includes a first prompt information and at least one data; the first prompt information is used to prompt the user to select data containing the first private entity; a fourth display module, used to display a fourth display interface in response to the user's selection operation in the at least one data, the fourth display interface includes first data and a first identifier, the first identifier is used to indicate that the first data is the data selected by the user.

24. The device according to claim 22 or 23, characterized in that The second display interface also includes: a second prompt message, the second prompt message is used to prompt the user to confirm whether the first search result is the search result expected by the user; the device also includes: a fifth display module, used to display a fifth display interface in response to the user's confirmation operation, the fifth display interface includes the first search result and a second identifier, the second identifier is used to indicate that the first search result is the search result confirmed by the user.

25. A model training device, characterized in that: The device includes: an acquisition module, used to obtain a multimodal sample set and a first model; wherein the multimodal sample set includes: multimodal samples corresponding to a first private entity, multimodal samples corresponding to a first search request, and multimodal samples corresponding to a second search request; the first private entity represents an entity having an association with a user, the first search request includes the first private entity, and the second search request does not include the first private entity; and a training module, used to train the first model using the multimodal sample set to obtain a second model.

26. The device according to claim 25, characterized in that The training module is also used to: process the multimodal samples corresponding to the first private entity and the multimodal samples corresponding to the first search request through the first model to obtain fused features; train the first model according to the fused features, the features corresponding to the first search request and the features corresponding to the second search request to obtain the second model; wherein the features corresponding to the first search request are obtained by processing the multimodal samples corresponding to the first search request by the first model, and the features corresponding to the second search request are obtained by processing the multimodal samples corresponding to the second search request by the first model.

27. The device according to claim 26, characterized in that The training module is also used to: input the multimodal sample corresponding to the first search request into the first model to generate a first feature; input the multimodal sample corresponding to the first private entity into the first model to generate a second feature; and fuse the first feature and the second feature to obtain the fused feature.

28. The device according to claim 26 or 27, characterized in that The multimodal samples include samples of the first modality and samples of the second modality; the features corresponding to the first search request are obtained by processing the samples of the first modality corresponding to the first search request by the first model, and the features corresponding to the second search request are obtained by processing the samples of the first modality corresponding to the second search request by the first model; the training module is also used to: input the samples of the first modality corresponding to the first private entity and the samples of the second modality corresponding to the first search request into the first model to obtain the fused features.

29. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the method of any one of claims 1-7, the method of claims 8-10, or the method of any one of claims 11-14 when executing the instructions.

30. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7, the method according to claims 8 to 10, or the method according to any one of claims 11 to 14 is implemented.