Retrieval method, device and system and terminal equipment
Patent Information
- Application Number
- CN202380094792.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2025-10-03
Smart Images

Figure CN120752627A_ABST
Abstract
Description
A search method, device, system and terminal device Technical Field
[0001] The present invention relates to the field of information technology (IT), and in particular to a retrieval method, apparatus, system and terminal equipment. Background Art
[0002] A search engine is a system that uses specific computer programs to collect information from the internet based on a specific strategy, organizes and processes the information, and then provides search services to users. Currently, mainstream search engines in the industry are cloud-based, enabling significant development of cloud-based search capabilities. For search operations, indexing technologies include traditional word-based inverted indexes and the currently popular approximate nearest neighbor (ANN) index based on model-based inference vectors. However, on the terminal side (such as mobile phones and tablets), resource limitations (such as memory, storage, CPU, and power consumption) prevent some excellent cloud-based search engines from being ported to the terminal side, resulting in poor search capabilities on these devices. Therefore, improving the search capabilities of terminal devices has become a pressing technical challenge.
[0003] Summary of the Invention
[0004] In order to solve the problems existing in the prior art, the embodiments of the present application provide a retrieval method, apparatus, system, terminal device, computer storage medium and product containing a computer program, which can borrow the resources of the server to enhance the retrieval capability of the terminal device.
[0005] In a first aspect, an embodiment of the present application provides a retrieval method, applied to a terminal device, comprising: obtaining a query statement input by a user; filtering out k target feature vectors from the vector index file based on the similarity between the query statement and target feature vectors stored in a vector index file, wherein the vector index file is obtained by processing intermediate feature vectors of objects obtained from the terminal device by a server and transmitted by the server to the terminal device, wherein the intermediate feature vectors are output results of a portion of a network layer in a neural network on the terminal device; and displaying objects corresponding to the K target feature vectors. For example, the objects may be, but are not limited to, images.
[0006] In this way, during the search process, the server and terminal device collaborate to generate a vector index file, allowing the search to find objects that match the user's query. Because the vector index file is jointly processed by the server and the terminal device, the object's feature vector is acquired by both the terminal device and the server. This reduces the workload on the terminal device and lowers the requirements for its capabilities, thereby leveraging server resources to enhance the terminal's search capabilities.
[0007] In one possible implementation, before selecting k target feature vectors from the vector index file based on the similarity between the query statement and the target feature vectors stored in the vector index file, the method further includes: processing the object in the terminal device using a first network layer in a neural network to obtain an intermediate feature vector of the object, wherein the neural network is obtained by the terminal device from a server and the neural network is mainly composed of a first network layer and a second network layer; sending a first message to the server, the first message being used to instruct the server to process the intermediate feature vector of the object using the second network layer in the neural network to obtain the vector index file; obtaining a second message from the server, the second message including the vector index file; and storing the vector index file in response to the second message. In this way, through the collaborative work of the terminal device and the server, the terminal device can leverage resources on the server to enhance its own retrieval capabilities.
[0008] In a second aspect, an embodiment of the present application provides a retrieval system, which includes a terminal device and a server, wherein the terminal device and the server are both configured with a neural network including N network layers, and the neural network on the terminal device is obtained from the server; wherein the terminal device is used to use the first M (M<N) network layers of the neural network to process at least one object to obtain an intermediate feature vector of the object, and to send a first message to the server, wherein the first message includes the intermediate feature vector of the object; the server is used to process the intermediate feature vector of the object using the M+1th to Nth network layers in the neural network model in response to the first message to obtain a target feature vector of the object; the server is also used to generate a vector index file based on the target feature vector of the object, wherein the vector index file is used to store the target feature vector; and, send a second message to the terminal device, wherein the second message includes the vector index file; the terminal device is also used to store the vector index file in response to the second message.
[0009] In this way, the terminal device and server together form a retrieval system. When a search is needed, the task of generating the object's feature vector can be jointly completed by the terminal device and the server. This can reduce the workload of the terminal device and reduce its resource usage, thereby improving the terminal device's retrieval capabilities by leveraging server resources.
[0010] In one possible implementation, the terminal device is further used to process the query statement input by the user through a neural network to obtain a first feature vector of the query statement; the terminal device is further used to filter out k target feature vectors from the vector index file based on the similarity between the first feature vector and the target feature vector stored in the vector index file, and to display the objects corresponding to the K target feature vectors.
[0011] In one possible implementation, the server is further used to: classify objects based on intermediate feature vectors of the objects sent by the terminal device to obtain the category and quantity of the objects; and, based on the category and quantity of the objects, calculate a personalized feature vector used to characterize the user of the terminal device; group the personalized feature vectors using a clustering algorithm to obtain a target cluster; use the center vector of the target cluster to filter out a personalized training data set related to the terminal device from a public training data set; use the public training data set to set the personalized training data set to update the neural network, and transmit the updated neural network to the terminal device.
[0012] In this way, by updating the neural network used by the terminal device based on the terminal device's own data, the neural network can be more in line with the usage habits of the user of the terminal device after training, and thus, when searching, it can better retrieve objects that conform to the user's habits.
[0013] In the third aspect, an embodiment of the present application provides a retrieval device, which is deployed on a terminal device and includes: an acquisition module for acquiring a query statement input by a user; a processing module for filtering out k target feature vectors from a vector index file based on the similarity between the query statement and the target feature vector stored in the vector index file, wherein the vector index file is obtained by the terminal device from a server and is obtained by the terminal device and the server collaboratively processing the objects in the terminal device; a display module for displaying the objects corresponding to the K target feature vectors.
[0014] In one possible implementation, before the processing module selects k target feature vectors from the vector index file based on the similarity between the query statement and the target feature vector stored in the vector index file, it is also used to: process the object in the terminal device through the first part of the network layer in the neural network to obtain the intermediate feature vector of the object, wherein the neural network is obtained by the terminal device from the server, and the neural network is mainly composed of the first part of the network layer and the second part of the network layer; send a first message to the server, the first message is used to instruct the server to use the second part of the network layer in the neural network to process the intermediate feature vector of the object to obtain the vector index file; obtain the second message sent by the server, the second message includes the vector index file; and store the vector index file in response to the second message.
[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium comprising computer-readable instructions. When a computer reads and executes the computer-readable instructions, the computer executes the method as described in any one of the first aspects.
[0016] In a fifth aspect, an embodiment of the present application provides a terminal device, comprising a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the method as described in any one of the first aspects is executed.
[0017] In a sixth aspect, an embodiment of the present application provides a product comprising a computer program, which, when the computer program product runs on a processor, enables the processor to execute the method as described in any one of the first aspects.
[0018] It can be understood that the beneficial effects of the third to sixth aspects mentioned above can be found in the relevant descriptions of the first and second aspects mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] FIG1 is a schematic diagram of the architecture of a terminal device provided in an embodiment of the present application;
[0021] FIG2 is a schematic diagram of a picture provided in an embodiment of the present application;
[0022] FIG3 is a schematic diagram of the architecture of another terminal device provided in an embodiment of the present application;
[0023] FIG4 is a schematic diagram of the architecture of a retrieval system provided in an embodiment of the present application;
[0024] FIG5 is a schematic diagram of an interface for opening a gallery provided in an embodiment of the present application;
[0025] FIG6 is a schematic diagram of a multi-modal reasoning engine provided in an embodiment of the present application;
[0026] FIG7 is a schematic diagram of a multi-modal reasoning engine training process provided by an embodiment of the present application;
[0027] FIG8 is a schematic diagram of a flow chart of a retrieval method provided in an embodiment of the present application;
[0028] FIG9 is a schematic diagram of a retrieval device provided in an embodiment of the present application;
[0029] FIG10 is a schematic diagram of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0031] The term "and / or" as used herein describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " as used herein indicates that the related objects are in an "or" relationship, for example, A / B means either A or B.
[0032] The terms "first" and "second" in this specification and claims are used to distinguish different objects rather than to describe a specific order of objects. For example, "first response message" and "second response message" are used to distinguish different response messages rather than to describe a specific order of response messages.
[0033] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0034] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.
[0035] To facilitate understanding of the embodiments of the present application, further explanation will be given below with reference to specific embodiments in conjunction with the accompanying drawings. The embodiments do not constitute a limitation on the embodiments of the present invention.
[0036] First, let’s introduce the technical terms involved in this application:
[0037] 1. Modality refers to the form in which data exists, such as text, images, videos, audio, etc.
[0038] 2. Modal retrieval is divided into unimodal retrieval and multimodal retrieval. Unimodal retrieval refers to searching using a query that is consistent with a particular modality, such as searching for images by image, searching for text by text, or searching for videos by video. Multimodal retrieval, also known as multimodal retrieval, is a technique for searching using queries that are inconsistent with a particular modality. It is a search between different categories, such as searching for images by text or searching for text by images.
[0039] 3. On-device multimodal search refers to the technology of performing multimodal search on mobile phones, tablets, PCs and other terminal devices.
[0040] 4. Cross-terminal multimodal retrieval refers to a technology that can distribute retrieval requests to multiple terminals through a certain data protocol, thereby realizing multimodal retrieval and result aggregation across multiple terminal devices.
[0041] 5. End-cloud collaboration refers to a mechanism that enhances end-side (mobile, generally referring to terminal devices) capabilities by using cloud-side (server-side) capabilities.
[0042] 6. Data desensitization refers to the deformation of certain sensitive information through desensitization rules to achieve reliable protection of sensitive privacy data. It is a technology that transforms real data without violating system rules.
[0043] Next, the technical solutions provided in the embodiments of the present application are introduced.
[0044] Generally, when performing data retrieval, indexing can improve retrieval efficiency. Therefore, when performing retrieval on a terminal device, to enhance the terminal's retrieval capabilities, it is common to tag the search content. The tags of the search content can be understood as indexes. Furthermore, to enhance search capabilities during retrieval, terminal devices generally have multimodal retrieval capabilities.
[0045] In this embodiment, the terminal device with multimodal retrieval capability may be configured with software and hardware. The hardware in the terminal device may provide an operating environment for the software in the terminal device.
[0046] For example, FIG1 shows a schematic diagram of the architecture of a terminal device provided by an embodiment of the present application. As shown in FIG1 , the terminal device includes software 100 and hardware 200. Software 100 may include a gallery 101, a tagging engine 102, a search engine 103, and an operating system (OS) 104. The gallery 101 may provide functions such as image viewing, a search portal, and display of search results. The tagging engine 102 may be composed of a neural network, which may generate at least one tag for each image in the gallery 101. For example, when an image in the gallery 101 is the image shown in FIG2 , the tags generated for the image by the tagging engine 102 may be "landscape" or "dog." The search engine 103 may be composed of a neural network, which may search the tags generated by the tagging engine 102 based on a query entered by a user in the search portal provided by the gallery 101 to obtain tags that match the user's query, and notify the gallery 101 to display images corresponding to tags that match the user's query. The OS 104 may provide basic resource scheduling management and access capabilities. Among them, the normal operation of the tag engine 102 and the retrieval engine 103 requires the support of OS 104, and the normal operation of OS 104 depends on hardware 200. Hardware 200 may include a central processing unit (CPU) 201, a memory chip 202, and a memory 203. Among them, CPU 201 may provide computing power to support the normal operation of the retrieval function. The memory chip 202 may provide a running space for the retrieval program, such as a random access memory (RAM). The memory 203 may provide persistence capabilities for the tags generated by the tag engine 102, such as a read-only memory (ROM) and a small storage (trans-flash card, TF) card, a secure digital memory card (SD) card, etc.
[0047] In the architecture of the terminal device shown in FIG1 , images are mainly retrieved based on their labels. In this retrieval method, it is often necessary to strictly match the keywords contained in the query input by the user with the pre-generated image labels, resulting in poor flexibility when performing retrieval. Since the label generated by the terminal device for the image is a scalar, it is impossible to effectively represent the semantics of the image. For example, referring to FIG2 , the labels generated by the terminal device for the image can only be "scenery" and "dog", but the query sentence input by the user is "black dog". In this case, the retrieval method cannot understand the user's retrieval request and cannot retrieve the image shown in FIG2 for the user.
[0048] Exemplarily, FIG3 shows a schematic diagram of the architecture of another terminal device provided in an embodiment of the present application. As shown in FIG3 , the terminal device includes software 310 and hardware 320, and the hardware 320 can provide an operating environment for the software 310. Among them, the software 310 includes a gallery 311, a multi-modal reasoning engine 312, a retrieval engine 313, and an OS 314. The gallery 311 can be the gallery 101 in FIG1 , and the OS 314 can be the OS 104 in FIG1 . Their functions can refer to the relevant description in FIG1 , and will not be repeated here. The multi-modal reasoning engine 312 can be composed of a neural network, which can extract feature vectors for different modalities. For example, it can extract feature vectors of pictures in the gallery, extract feature vectors of text (query) entered by the user in the retrieval entrance in the gallery, and so on. In addition, the multi-modal reasoning engine 312 can also transmit the feature vector extracted by the query to the retrieval engine 313. The retrieval engine 313 can be composed of a neural network, which searches the feature vectors of the pictures in the gallery generated by the inference engine 312 based on the feature vector of the query transmitted by the inference engine 312 to obtain pictures that match the feature vector of the query, and notifies the gallery 311 to display pictures that match the feature vector of the query.
[0049] Hardware 320 includes a CPU 321, a memory chip 322, a memory 323, an embedded neural network processing unit (NPU) 324, and a graphics processing unit (GPU) 325. CPU 321 may be CPU 201 in Figure 1 , the memory chip may be memory chip 202 in Figure 1 , and memory 323 may be memory 203 in Figure 1 . Their functions can be referred to in the relevant descriptions in Figure 1 and will not be repeated here. NPU 324 and GPU 325 can provide the computing power required for machine learning and neural network operators to run programs.
[0050] In the terminal device architecture shown in Figure 3, image retrieval is primarily based on the image's feature vector. This retrieval method provides a comprehensive representation of images from a semantic perspective, enabling the search engine to identify the user's semantic information and meet their diverse search needs. However, this method places high demands on the terminal device's computing power and resources such as random access memory (RAM). Physical limitations (such as size) limit the resources available on terminal devices. When processing complex machine learning model calculations or inference processes, excessive resource usage can significantly impact system resource management. Furthermore, using the same multimodal inference engine on different terminal devices prevents personalized information (i.e., tailored to the user's specific needs) from being provided to different users. For example, if different users enter the same query "beetle," a car enthusiast would expect to retrieve images containing beetle-shaped cars, while an insect enthusiast would prefer to retrieve images containing beetles.
[0051] In view of this, an embodiment of the present application provides a retrieval method that primarily infers the feature vector of an object to be retrieved through a terminal device and a server. This method divides the task that would otherwise be completed independently by the terminal device into two parts: one completed on the terminal device and the other completed on the server. In this way, the inference of the object's feature vector can be completed jointly by the terminal device and the server. This significantly reduces the workload of the terminal device, which in turn reduces its resource usage, thereby enhancing the terminal device's retrieval capabilities by utilizing server resources.
[0052] Exemplarily, FIG4 shows a schematic diagram of the architecture of a retrieval system provided by an embodiment of the present application. FIG4 is an example of an image being the object to be retrieved. As shown in FIG4 , the image retrieval system 400 may include a terminal device 410 and a server 420. The terminal device 410 and the server 420 may communicate via a wired network or a wireless network. The terminal device 410 may be a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook or a mobile phone, etc. The embodiment of the present application does not impose any special restrictions on the specific type of the terminal device 410. The server 420 may be a central server, an edge server, a cloud server, or a local server in a local data center, etc. The embodiment of the present application does not impose any special restrictions on the specific type of the server 420. Exemplarily, when the server 420 is a cloud server, users can use the services provided by the server 420 by purchasing the services provided by the server 420.
[0053] The terminal device 410 may include a gallery 411, a multimodal reasoning engine 412, a terminal-cloud collaboration framework 413, and a vector engine 414. The gallery 411 may be used to display images or videos stored in the terminal device 410. The images or videos stored in the terminal device 410 may be captured by the user using a camera (e.g., a camera) or downloaded via other applications (e.g., a browser). Corresponding identifiers may be "album," "photo," or the like. The gallery 411 may have different controls, and when the user selects different controls, the gallery 411 may display different interfaces. After entering the gallery, different controls may be displayed on the gallery interface, and each interface displayed by the gallery 411 may include a search entry for the user to enter a query. For example, as shown in FIG5(A), the display interface of the terminal device 410 displays multiple apps, including the gallery 411, and the user may click on the gallery 411 with a finger 51. Afterwards, the terminal device 410 may display an interface as shown in (B) of FIG5 . In the interface shown in (B) of FIG5 , a control 52 for the user to input a query may be displayed. In addition, as shown in (B) of FIG5 , the gallery 411 may also display controls for classifying images into different categories, such as “Photos,” “Albums,” “Moments,” and “Discovery.”
[0054] The multimodal reasoning engine 412 can be a neural network (NN), such as a convolutional neural network (CNN) or a deep neural network (DNN). The multimodal reasoning engine 412 is mainly used to extract features of the query input by the user in the gallery 411 to obtain the feature vector of the query, and to extract features of the pictures in the gallery 411 through a part of the network layer therein to obtain the intermediate feature vectors of each picture in the gallery 411. Exemplarily, as shown in Figure 6, the multimodal reasoning engine 412 can include N (N≥2) network layers. These N network layers are divided into two parts by the bottleneck layer, namely the base part consisting of the first M layers (M<N) and the bottleneck layer, and the delta part consisting of the bottleneck layer and the network layers between the M+1th layer and the Nth layer. Among them, the bottleneck layer refers to the middle layer between the input layer and the output layer in the neural network, which uses a 1*1 convolutional neural network, and its output dimension is much smaller than the input dimension. The bottleneck layer is usually used for dimensionality reduction and feature extraction, which can effectively reduce network parameters and computational complexity, improve the training speed and generalization ability of the model, and in common deep neural networks such as ResNet and Inception, the bottleneck layer design is adopted, which can make the network more efficient and accurate. When determining the bottleneck layer, the bottleneck layer can be searched through the neural architecture search (NAS) technology. When the multi-modal reasoning engine 412 in the terminal device 410 processes the image, the base part of the multi-modal reasoning engine 412 can be used. In some embodiments, the intermediate feature vector of the image can be, but is not limited to, the desensitized feature data obtained after the image is processed by the base part of the multi-modal reasoning engine 412, wherein the desensitized feature data refers to the data obtained after the data desensitization operation. If the desensitized feature data is restored and the original data cannot be obtained, it can be considered that the data desensitization is successful. Of course, in some embodiments, with the user's permission (for example, the user agrees to the user agreement, and it is agreed in the agreement that all pictures or videos of the user can be used), the user's data may not be desensitized. It is worth noting that due to the different OS of the terminal device, the multi-mode reasoning engine 412 may also be different for different OS. For example, if the terminal device is an OS based on the Hongmeng system, the multi-mode reasoning engine 412 is a multi-mode reasoning engine adapted to the Hongmeng system; if the terminal device is an OS based on the Apple system, the multi-mode reasoning engine 412 is a multi-mode reasoning engine adapted to the Apple system.
[0055] For example, when the multimodal inference engine 412 extracts features from an image, it can utilize its base portion to extract features from the image. When the multimodal inference engine 412 extracts features from a user-input query, it can utilize its N network layers to extract features from the query. In a digital image, each pixel represents an optical point and contains information such as the color and brightness of that point. Typically, each pixel in an image consists of three color channels (red, green, and blue, RGB), each of which can have a value between 0 and 255. Therefore, for an image, each pixel contains a wealth of information, and the image content composed of these pixels can be very large. If all image processing were performed on the terminal device by the multimodal inference engine 412 of the terminal device, the terminal device would consume a significant amount of resources, impacting the user's normal use of the terminal device. Therefore, it is possible to choose to use only the base portion of the multimodal inference engine 412 to extract features from the image, thereby occupying a small amount of terminal device resources. Since textual content contains less information, the N network layers of the multimodal reasoning engine 412 can be used to extract features from the query input by the user, without occupying a large amount of terminal device resources. For example, when the user first launches the software (such as the gallery 411), the terminal device 410 sends a multimodal reasoning engine acquisition request to the server 420 to obtain the multimodal reasoning engine 412 from the server 420. Afterwards, the terminal device 410 can also obtain the multimodal reasoning engine 412 from the server 420 in real time or periodically (for example, every 1 day, 15 days, or 1 month, etc.). Alternatively, the terminal device 410 can send a query request to the server 420 each time it is opened or at a preset time, requesting the version number of the multimodal reasoning engine 412. When the queried version number is different from the version number of the local multimodal reasoning engine 412, the multimodal reasoning engine 412 is obtained from the server 420.
[0056] The end-cloud collaboration framework 413 is mainly used to provide an end-cloud data communication and transmission framework directly or indirectly constructed based on the OS of the terminal device 410, so that the terminal device 410 has the ability to work in collaboration with the server 420. Among them, the terminal device 410 can periodically send the intermediate vector of the image to the server, for example, every 5 days, 10 days, etc., or the terminal device 410 and the server 420 can determine the time when the terminal device 410 sends the intermediate vector of the image to the server 420 through negotiation. Among them, the first feature vector of the image extracted by the multi-modal reasoning engine 412 can be transmitted to the server 420 through the end-cloud collaboration framework 413, so that the server 420 can further perform feature extraction on each image based on the intermediate feature vector of each image to obtain the target feature vector of each image, that is, the complete feature vector. Exemplarily, the end-cloud collaboration framework 413 can be the function flow runtime (FFRT) in Huawei mobile service (HMS).
[0057] The vector retrieval engine 413 is mainly used to provide vector retrieval capabilities. Through the vector retrieval engine 413, the target feature vector of the image extracted by the server 420 obtained through the end-cloud collaborative framework 413 can be persisted, and, based on the feature vector of the user-input query extracted by the multi-modal reasoning engine 412, the image that matches the feature vector of the user-input query is retrieved, and the retrieved image is fed back to the gallery 411 so that the gallery 411 displays these images. Exemplarily, after the vector retrieval engine 413 obtains the feature vector of the user-input query extracted by the multi-modal reasoning engine 412, it can calculate the similarity between the feature vector and the target feature vector of each image in the gallery 411, and use the images corresponding to the top K target feature vectors with higher similarity to the feature vector of the user-input query as the retrieved images.
[0058] The server 420 may include: an end-cloud collaboration framework 421, a multi-modal reasoning engine 422, and a vector indexing engine 423. Among them, the end-cloud collaboration framework 413 is mainly used to provide an end-cloud data communication and transmission framework directly or indirectly constructed based on the OS of the server 420, so that the server 420 has the ability to work in collaboration with the terminal device 410. Among them, the server 420 can obtain the first feature vector of the image extracted by the multi-modal reasoning engine 412 in the terminal device 410 through the end-cloud collaboration framework 421, and obtain the identification of the terminal device 410. The server 420 can also transmit the feature vector extracted by the multi-modal reasoning engine 422 to the terminal device 410 through the end-cloud collaboration framework 421. Exemplarily, the end-cloud collaboration framework 421 can be communications-as-a-service (CaaS) in the HMS.
[0059] The multimodal reasoning engine 422 is the same as the multimodal reasoning engine 412. The multimodal reasoning engine 422 can further perform feature extraction on the intermediate feature vector of the image extracted by the multimodal reasoning engine 412 on the terminal device 410 to obtain the target feature vector of the image. Exemplarily, when the multimodal reasoning engine 422 is composed of N network layers, when the multimodal reasoning engine 412 uses the first M (M<N) network layers therein to extract features from the image, the multimodal reasoning engine 422 can use the delta part therein to perform feature extraction on the intermediate feature vector of the image extracted by the multimodal reasoning engine 412 to obtain the complete feature vector of the image, that is, the target feature vector. The vector extraction process can be completed on a cloud computing platform, and the cloud computing platform provides the computing power for vector extraction. The multimodal reasoning engine 422 can be updated periodically, for example, the multimodal reasoning engine 422 is trained once every 10 days. A user agreement exists in the terminal device 410, which stipulates that the server 420 can obtain the intermediate feature vectors of the images in the user's gallery at a certain time. For example, every 10 days, between 2:00 and 4:00 in the evening, the terminal device 410 sends the intermediate feature vectors of the images in the gallery to the server 420. If the user agrees to this agreement, when the multimodal inference engine 422 is retrained, the complete feature vectors of the images in the gallery 411 of the terminal device 410 can be obtained based on the intermediate feature vectors of the images sent by the terminal device 410, and the multimodal inference engine 422 can be retrained based on the obtained complete feature vectors. After the multimodal inference engine 422 is retrained, the retrained multimodal inference engine 422 is stored, replacing the previous version of the multimodal inference engine 422. When the multimodal inference engine 422 is stored, the version number of the multimodal inference engine 422 can also be stored.
[0060] The vector index engine 423 is primarily used to generate a vector index file based on the target feature vector extracted by the multimodal reasoning engine 422. This vector index file can be used to store the target feature vector. The vector index file can be in a .tvx, .tvd, or .tvf format, for example. Exemplarily, the vector index file may include an index structure and the target feature vector of at least one image in the image library 411. The index structure refers to a data structure used to organize and accelerate vector retrieval, such as a KD-Tree, LSH, or B-Tree.
[0061] As can be seen from the above, in the embodiment of the present application, when obtaining the high-dimensional spatial feature vector of an image, the feature vector reasoning process is divided into two parts: one part is completed by the base part of the multimodal reasoning engine on the terminal device, and the other part is completed by the delta part of the multimodal reasoning engine on the server. In this way, the complete feature vector of the image is obtained through end-cloud collaboration. This can significantly reduce the power consumption of the terminal device, and also reduce the resource usage of the terminal device, creating the premise for the commercialization of multimodal semantic search on terminal devices.
[0062] To enable different multimodal reasoning engines for different users, thereby satisfying their personalized requests during searches, the multimodal reasoning engines can also be trained individually. The personalized training of the multimodal reasoning engines can be performed on a cloud computing platform, which provides model training capabilities. For example, the multimodal reasoning engine can be multimodal reasoning engine 422 in Figure 4 .
[0063] Next, the training process of the multi-modal reasoning engine 422 is introduced based on the content of FIG. 4 .
[0064] For example, Figure 7 shows a schematic diagram of a multi-modal reasoning engine training process provided by an embodiment of the present application. As shown in Figure 7, the process of training the multi-modal reasoning engine may include the following steps:
[0065] S701: Classify the intermediate feature vectors sent by each terminal device to obtain the category and number of objects corresponding to each intermediate feature vector, and calculate the personalized feature vectors for representing the users of the terminal devices based on the category and number of each object.
[0066] In this embodiment, after server 420 receives the user's data (the intermediate feature vector of the image), the multimodal inference engine 422 on server 420 can perform preliminary processing on the intermediate feature vector of the image to obtain the number of images in each category in the image gallery 411 of the user terminal device 410. For example, it can be determined that there are 100 images in the scenery category and 50 images in the food category in the user's gallery. The obtained categories and numbers are then processed using a bag-of-words model to obtain a bag-of-words vector. The bag-of-words model is a text representation method that views text as a collection of words, ignoring the order and grammar of the words, as well as the relationships between them, and focusing only on the frequency of word occurrence in the text. Specifically, the bag-of-words model represents text as a vector, with each dimension of the vector corresponding to a word, and the value of each dimension representing the frequency of the word's occurrence in the text. For example, a vector [100, 50] obtained after processing by the bag-of-words model indicates that the scenery category appears 100 times and the food category appears 50 times. In this way, a single text can be represented by a single vector, and multiple texts can be composed into a text matrix. For example, suppose there are two texts as follows:
[0067] Text 1: I like machine learning.
[0068] Text 2: I like deep learning.
[0069] Considering only the probability of word occurrence in a text, let's say that text 1 contains the words "I," "like," "machine," and "learning," and text 2 contains the words "I," "like," "deep," and "learning." That is, text 1 and text 2 share the five words "I," "like," "machine," "learning," and "deep." Therefore, text 1 can be represented by the vector [1,1,1,1,0], where the first dimension has a value of 1, indicating that the word "I" appears once in text 1, and the last dimension has a value of 0, indicating that the word "deep" does not appear in text 1. Similarly, text 2 can be represented by the vector [1,1,0,1,1]. After obtaining the bag-of-words vector, it can be normalized. In machine learning, data is typically represented as vectors, where each dimension corresponds to a different feature (i.e., a category, such as scenery or food). When processing high-dimensional vectors, normalization is often necessary to ensure that each feature contributes equally to the model's training results. Normalization is a common data preprocessing method that aims to map data features of varying scales to the same scale for better comparison and analysis. Through normalization, a unique feature vector for a single user (i.e., a single terminal device 410) can be obtained, which is known as a personalized feature vector for that user.
[0070] S702: Grouping the personalized feature vectors for characterizing the users of the terminal device by using a clustering algorithm to obtain at least one cluster, wherein each cluster includes at least one personalized feature vector.
[0071] In this embodiment, server 420 obtains a personalized feature vector for each user. Server 420 can then cluster the personalized feature vectors of a large number of users using a clustering model. Clustering is an unsupervised learning technique used to divide objects in a dataset into several categories or clusters. The goal of clustering is to maximize the similarity between objects within a category and minimize the similarity between different categories. Clustering algorithms typically calculate based on the similarity or distance between data objects. Common clustering algorithms include K-means clustering, hierarchical clustering, and DBSCAN. Applications of clustering algorithms include image segmentation, text mining, and social network analysis. The advantage of clustering algorithms is that they can automatically discover the inherent structure and patterns in a dataset without pre-setting category labels, making them suitable for processing large datasets. After clustering, multiple user groups can be obtained, and each of the multiple cluster centers obtained by clustering is considered a user group. A cluster center is a vector, meaning that a user group can be represented by a vector. Each user corresponds to a user group.
[0072] S703: Based on the vectors of the centers of the clusters, a personalized training set corresponding to each cluster is selected from the public training set.
[0073] In this embodiment, the server 420 pre-prepares a public training set, which can be classified by category to obtain public training sets of different categories, denoted as public training set 1, public training set 2, ..., public training set n, where public training set 1 can be classified as scenery, public training set 2 can be classified as food, and so on. In this way, the initial public training set can be classified into n public training sets, each of which stores image data of the corresponding category. Next, the n public training sets can be sampled according to the proportion of each dimension in the vectors of different cluster centers obtained in the clustering to obtain feature training sets for different user groups. For example, the vector of a user group is represented as [0.15, 0.3, 0.25, 0.2], where the first dimension represents scenery, the second dimension represents food, the third dimension represents cars, and the fourth dimension represents portraits. For example, if 100 images are required in each feature training set, 15 images are selected from the public training set classified as scenery, 30 images are selected from the public training set classified as food, 25 images are selected from the public training set classified as cars, and 20 images are selected from the public training set classified as portraits. The selected image data are integrated to form a training set corresponding to a single group, that is, a personalized training set for each group.
[0074] S704: Based on the public training set and each personalized training set, the multimodal reasoning engine used by the terminal devices in the group corresponding to the corresponding personalized training set is updated to obtain the personalized multimodal reasoning engine used by each terminal device.
[0075] In this embodiment, after obtaining a feature training set for each group, server 420 can fine-tune the multimodal inference engine 422 to obtain feature-based multimodal inference engines for different groups. Since the multimodal inference engine is trained using the initial public training set, the resulting multimodal inference engine can recognize the semantics of images, but cannot understand the different semantics of the same description. For example, it can recognize images of cars that resemble "beetles" and images of insects that belong to the same species, but cannot determine whether the user's query containing "beetles" actually wants to see pictures of cars or insects. Therefore, when fine-tuning the model, the training set can be the personalized training set for each group obtained in (C) of Figure 7 and the public training set. In this way, feature-based multimodal inference engines can be obtained for different user groups, thereby obtaining a multimodal inference engine 422 that can be targeted to different users.
[0076] Furthermore, server 420 can transmit the multimodal reasoning engine required by each terminal device to the corresponding terminal device. Afterwards, the terminal device can use the latest multimodal reasoning engine to reprocess the data within it. Because the latest multimodal reasoning engine is trained based on the personalized characteristics of the user of the terminal device (user), the data processed by the multimodal reasoning engine is more closely matched to the user of the terminal device and can better reflect the user's usage habits. Therefore, when performing subsequent searches, objects that are more consistent with the user's usage habits can be retrieved.
[0077] The above is an introduction to the architecture of the retrieval system provided in the embodiment of the present application. The following is an introduction to the method for performing retrieval using this architecture. For example, FIG8 shows a flow chart of a retrieval method provided in the embodiment of the present application. The method can be applied to a terminal device, such as the terminal device 410 in FIG4 . As shown in FIG8 , the retrieval method may include the following steps:
[0078] S801: Obtain the query statement input by the user.
[0079] In this embodiment, a user can enter a query statement in a search entry provided by a terminal device. After the user completes the input, the terminal device can retrieve the query statement entered by the user. After receiving the query statement entered by the user, the terminal device can process the query statement through the multimodal reasoning engine on the terminal device to obtain a feature vector for the query statement.
[0080] S802: Based on the similarity between the query statement and the target feature vector stored in the vector index file, k target feature vectors are filtered out from the vector index file, wherein the vector index file is obtained by the server by processing the intermediate feature vector of the object obtained from the terminal device, and is transmitted to the terminal device by the server, and the intermediate feature vector is the output result of a part of the network layer in the neural network on the terminal device.
[0081] In this embodiment, after obtaining the feature vector of the query statement entered by the user, the terminal device may first calculate the similarity between the feature vector of the query statement and the target feature vector of the object stored in the vector index file. The terminal device may then filter K target feature vectors from the vector index file based on the calculated similarity. The K selected target feature vectors may be the K values with the highest similarity or the K values whose similarity exceeds a set threshold. For example, the object may be, but is not limited to, an image.
[0082] In addition, the vector index file can be obtained by the terminal device from the server. The process of obtaining the vector index file can be: the terminal device first processes the object (such as: picture, etc.) in it through the first part of the network layer in the neural network therein to obtain the intermediate feature vector of the corresponding object. Afterwards, the terminal device can transmit the intermediate feature vector it obtained to the server. For example, the terminal device can send a message including the intermediate feature vector to the server, and the message can be used to instruct the server to process these intermediate feature vectors using the second part of the network layer in the neural network to obtain the vector index file. After obtaining the intermediate feature vector, the server can use the remaining network layers in the neural network to process the intermediate feature vector to obtain the vector index file. Finally, the server can transmit the vector index file to the terminal device. After obtaining the vector index file, the terminal device can store the vector index file. Exemplarily, the neural network in the terminal device and the neural network in the server are the same, and the neural network in the terminal device is obtained from the server.
[0083] S803: Display objects corresponding to the K target feature vectors.
[0084] In this embodiment, after obtaining K target feature vectors, the terminal device can display the objects corresponding to the K target feature vectors on its display component. For example, the terminal device can display the pictures corresponding to the K target feature vectors on its gallery interface for the user to view.
[0085] During the search process, the server and the terminal device collaboratively process the vector index file, which can be used to search for objects that match the user's query. Since the vector index file is not processed independently by the terminal device, the terminal device's workload is reduced, which significantly reduces the terminal device's resource usage, thereby improving the terminal device's search capabilities by leveraging server resources.
[0086] It is understandable that the size of the sequence number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. In addition, in some possible implementations, the steps in the above embodiment can be selectively executed according to actual conditions, and can be partially executed or fully executed, which is not limited here. All or part of any features of any embodiment of the present application can be freely and arbitrarily combined without contradiction. The combined technical solution is also within the scope of the present application.
[0087] Based on the method in the above embodiment, an embodiment of the present application also provides a retrieval device.
[0088] 9 shows a retrieval apparatus, which may be deployed in the aforementioned terminal device 410. As shown in FIG9 , the retrieval apparatus 900 may include: an acquisition module 901, a processing module 902, and a display module 903.
[0089] The acquisition module 901 may be used to acquire a query statement input by a user.
[0090] The processing module 902 can be used to filter out k target feature vectors from the vector index file based on the similarity between the query statement and the target feature vectors stored in the vector index file, wherein the vector index file is obtained by the terminal device from the server, and is obtained by the terminal device and the server collaboratively processing the objects in the terminal device.
[0091] The display module 903 may be used to display objects corresponding to the K target feature vectors.
[0092] In some embodiments, before the processing module selects k target feature vectors from the vector index file based on the similarity between the query statement and the target feature vectors stored in the vector index file, the processing module 902 is further configured to:
[0093] An object in a terminal device is processed by a first network layer in a neural network to obtain an intermediate feature vector of the object, wherein the neural network is obtained by the terminal device from a server, and the neural network is mainly composed of a first network layer and a second network layer; a first message is sent to the server, and the first message is used to instruct the server to use the second network layer in the neural network to process the intermediate feature vector of the object to obtain a vector index file; a second message is sent by the server, and the second message includes the vector index file; and the vector index file is stored in response to the second message.
[0094] It should be understood that the above-mentioned device is used to execute the method in the above-mentioned embodiment. The implementation principle and technical effect of the corresponding program module in the device are similar to those described in the above-mentioned method. The working process of the device can refer to the corresponding process in the above-mentioned method and will not be repeated here.
[0095] Based on the methods in the above embodiments, embodiments of the present application further provide a terminal device 1000. As shown in Figure 10, terminal device 1000 includes a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. Processor 1004, memory 1006, and communication interface 1008 communicate with each other via bus 1002. It should be understood that this application does not limit the number of processors and memories in terminal device 1000. For example, terminal device 1000 may be terminal device 410 described above.
[0096] Bus 1002 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG10 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 1002 may include a path for transmitting information between various components of electronic device 1000 (e.g., memory 1006, processor 1004, and communication interface 1008).
[0097] The processor 1004 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0098] The memory 1006 may include volatile memory, such as random access memory (RAM). The processor 1004 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0099] Memory 1006 stores executable program code, which processor 1004 executes to implement the functions of the modules described above in FIG. 9 , thereby performing all or part of the steps of the method in the above-described embodiment. In other words, memory 1006 stores instructions for executing all or part of the steps of the method in the above-described embodiment.
[0100] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the terminal device 1000 and other devices or a communication network.
[0101] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.
[0102] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the method in the above embodiment.
[0103] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0104] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0105] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).
[0106] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.
Claims
1. A search method, characterized in that: Applied to terminal equipment, including: Get the query statement entered by the user; Based on the similarity between the query statement and the target feature vector stored in the vector index file, k target feature vectors are screened out from the vector index file, wherein the vector index file is obtained by processing the intermediate feature vector of the object obtained from the terminal device by the server and transmitted to the terminal device by the server, and the intermediate feature vector is the output result of a part of the network layer in the neural network on the terminal device; Display the objects corresponding to the K target feature vectors.
2. The method according to claim 1, characterized in that Before selecting k target feature vectors from the vector index file based on the similarity between the query statement and the target feature vectors stored in the vector index file, the method further includes: Processing the object in the terminal device through a first partial network layer in a neural network to obtain an intermediate feature vector of the object, wherein the neural network is obtained by the terminal device from the server, and the neural network is mainly composed of the first partial network layer and the second partial network layer; Sending a first message to the server, wherein the first message is used to instruct the server to process the intermediate feature vector of the object using the second part of the network layer in the neural network to obtain the vector index file; Obtaining a second message sent by the server, where the second message includes the vector index file; In response to the second message, the vector index file is stored.
3. A retrieval system, characterized in that: include: A terminal device and a server, wherein the terminal device and the server are both configured with a neural network including N network layers, and the neural network on the terminal device is obtained from the server; The terminal device is used to process at least one object using the first M (M<N) network layers of the neural network to obtain an intermediate feature vector of the object, and to send a first message to the server, wherein the first message includes the intermediate feature vector of the object; The server is configured to process the intermediate feature vector of the object using the M+1th to Nth network layers in the neural network model in response to the first message to obtain the target feature vector of the object; The server is further configured to generate a vector index file based on the target feature vector of the object, wherein the vector index file is used to store the target feature vector; and send a second message to the terminal device, wherein the second message includes the vector index file; The terminal device is further configured to store the vector index file in response to the second message.
4. The retrieval system according to claim 3, characterized in that: The terminal device is further used to process the query statement input by the user through the neural network to obtain a first feature vector of the query statement; The terminal device is further used to filter out k target feature vectors from the vector index file based on the similarity between the first feature vector and the target feature vector stored in the vector index file, and to display the objects corresponding to the K target feature vectors.
5. The retrieval system according to claim 3 or 4, characterized in that: The server is also used for: Based on the intermediate feature vectors of the objects sent by the terminal device, the objects are classified to obtain the categories and quantities of the objects; and based on the categories and quantities of the objects, a personalized feature vector for characterizing a user of the terminal device is calculated; Using a clustering algorithm to group the personalized feature vectors to obtain a target cluster; Using the vector of the center of the target cluster, a personalized training data set related to the terminal device is screened out from a public training data set; The neural network is updated using the public training data set and the personalized training data set, and the updated neural network is transmitted to the terminal device.
6. A search device, characterized in that: Deployed on terminal devices, including: The acquisition module is used to obtain the query statement input by the user; a processing module, configured to filter out k target feature vectors from the vector index file based on the similarity between the query statement and the target feature vector stored in the vector index file, wherein the vector index file is obtained by the terminal device from a server and is obtained by the terminal device and the server collaboratively processing an object in the terminal device, A display module is used to display objects corresponding to the K target feature vectors.
7. The search device according to claim 6, characterized in that: Before the processing module selects k target feature vectors from the vector index file based on the similarity between the query statement and the target feature vectors stored in the vector index file, the processing module is further configured to: Processing the object in the terminal device through a first partial network layer in a neural network to obtain an intermediate feature vector of the object, wherein the neural network is obtained by the terminal device from the server, and the neural network is mainly composed of the first partial network layer and the second partial network layer; Sending a first message to the server, wherein the first message is used to instruct the server to process the intermediate feature vector of the object using the second part of the network layer in the neural network to obtain the vector index file; Obtaining a second message sent by the server, where the second message includes the vector index file; In response to the second message, the vector index file is stored.
8. A terminal device, characterized in that: include: at least one memory for storing a program; at least one processor, configured to execute the program stored in the memory; When the program stored in the memory is executed, the processor is used to execute the method as claimed in claim 1 or 2.
9. A computer-readable storage medium storing a computer program, wherein when the computer program is executed on a processor, the processor is caused to execute the method according to claim 1 or 2.
10. A computer program product, characterized in that When the computer program product is run on a processor, the processor is caused to perform the method according to claim 1 or 2.