Data retrieval and management method based on terminal model
By using machine learning models on the terminal for data retrieval and management, the problem of difficulty in retrieving and managing multimodal cached data in the prior art is solved, and efficient data retrieval and convenient data management are achieved.
Patent Information
- Application Number
- CN202510215972.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-20
AI Technical Summary
It is difficult for the prior art to effectively retrieve and manage multimodal cached data, and users cannot quickly find the required cached data.
Using a terminal model-based data retrieval and management method, a machine learning model is used to identify and match the search content input by the user and the cached data of the terminal, providing an interface for the user to manage the retrieved data.
Improves the efficiency of data retrieval and management convenience, allowing users to quickly find and manage the required multimodal cached data.
Smart Images

Figure CN120179689A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data retrieval, and in particular to a data retrieval and management method based on a terminal model. Background Art
[0002] In the field of mobile Internet, applications are widely used in users' daily lives. Users can browse news, shop, and be entertained through applications. During the use of applications, a large amount of cached data is generated, including various types of cached data such as pictures, videos, audios, and texts. These cached data are only used when the program is loaded again and cannot be directly used by users.
[0003] In the related art, there is a cache cleaning entry in the application for users to clean the cache area, but users cannot effectively clean individual cache files, or there is a cache file retrieval entry in the application for users to retrieve cache files by characters. However, the characters used for retrieval can only match the cached text data and cannot match the files of other modal data such as cached pictures, videos, or voices. Summary of the Invention
[0004] To solve the deficiencies of the prior art, the purpose of this application is to provide a data retrieval and management method based on a terminal model. This method can retrieve multi-modal cached data through a machine learning model, enabling users to quickly find the required cached data and display it for users to manage, improving the retrieval efficiency and management convenience.
[0005] To achieve the above purpose, this application adopts the following technical solutions:
[0006] In a first aspect, this application provides a data retrieval and management method based on a terminal model, and the method includes:
[0007] Provide a first interface and obtain the retrieval content input by the user on the first interface; wherein, the retrieval content includes a retrieval object defined by the user.
[0008] Load the machine learning model stored in the terminal, and use the machine learning model to access the cached data of the terminal and determine whether the retrieval object exists in the cached data; wherein, the cached data is the data used when running the application on the terminal, and the used data includes at least one of the following: the data loaded by the application cache, the pre-built data, and the data downloaded in response to the user's download instruction.
[0009] Provide a second interface, on which an optional list of target data and buttons for managing at least some of the target data selected by the user from the optional list are displayed. The target data is the cached data containing the retrieval object retrieved by the machine learning model.
[0010] In response to a user's operation on a key, manage at least some of the selected target data; wherein, managing at least some of the selected target data includes one of the following: deleting at least some of the selected target data, downloading at least some of the selected target data.
[0011] In one embodiment, the machine learning model includes an identification model and a matching model. Using the machine learning model to access the cached data of the terminal and determine whether a retrieval object exists in the cached data includes:
[0012] Using the identification model to identify the retrieval content and the cached data, and using the matching model to match the identified results to determine whether a retrieval object exists in the cached data.
[0013] In one embodiment, the identification model includes a first identification model for identifying the retrieval content and a second identification model for identifying the cached data;
[0014] Wherein, if the format of the retrieval object is one or more of pictures, videos, and audios, the first identification model and the matching model are the same model.
[0015] In one embodiment, the second identification model includes an audio model for audio recognition and / or an image model for image recognition;
[0016] If the format of the retrieval object is pictures and / or videos, using the identification model to identify the retrieval content and the cached data, and using the matching model to match the identified results to determine whether a retrieval object exists in the cached data includes:
[0017] Encoding the retrieval content using the text encoder of the first identification model to extract semantic features, obtaining a first feature vector, and encoding the pictures and videos in the cached data using the picture encoder of the image model to extract visual features, obtaining multiple second feature vectors;
[0018] Using the matching model to compare the first feature vector with the multiple second feature vectors respectively, determining the matching degrees of the first feature vector with the multiple second feature vectors respectively, and determining whether the pictures and videos in the cached data contain the retrieval object according to the matching degrees.
[0019] In one embodiment, if the format of the retrieval object is audio, using the identification model to identify the retrieval content and the cached data, and using the matching model to match the identified results to determine whether a retrieval object exists in the cached data includes:
[0020] Encoding the retrieval content using the first identification model to extract semantic features, obtaining a third feature vector, and encoding the voice files in the cached data using the audio model to extract audio features, obtaining multiple fourth feature vectors;
[0021] Use a matching model to compare the third feature vector with multiple fourth feature vectors respectively, determine the matching degrees of the third feature vector with multiple fourth feature vectors respectively, and determine whether the voice files in the cached data contain the retrieval object according to the magnitudes of the matching degrees.
[0022] In one embodiment, using a matching model to compare the first feature vector with multiple second feature vectors respectively, and determining the matching degrees of the first feature vector with multiple second feature vectors respectively, includes:
[0023] Using the matching model, for any second feature vector, based on the second feature vector and the first feature vector, determine a matching degree matrix; wherein, the first feature vector includes T i text features, and any second feature vector includes I j image features, the rows of the matching degree matrix correspond to different text features T i and the columns of the matching degree matrix correspond to different image features I j and each element T i I j in the matching degree matrix is the matching degree between the corresponding image feature and the corresponding text feature;
[0024] Using the matching model, based on each matching degree matrix, determine the matching degrees of the first feature vector with each second feature vector respectively.
[0025] In one embodiment, the data retrieval and management method further includes: obtaining the attributes of the cached data and / or the attributes of the application environment, and obtaining a machine learning model from the cloud according to the attributes of the cached data and / or the attributes of the application environment; wherein, multiple machine learning models with different parameter amounts are stored in the cloud.
[0026] In one embodiment, the data retrieval and management method further includes: providing a third interface, obtaining the set conditions and management operations for the cached data input by the user on the third interface, wherein when the cached data meets the set conditions, perform the management operations for the cached data.
[0027] In one embodiment, the set conditions at least include one or more of the cached data exceeding a set quantity, the generation time exceeding a timeout time, the access time exceeding a set time, and the access times exceeding a set number of times; the management operations at least include deletion and export to a preset storage location.
[0028] In one embodiment, loading the machine learning model stored in the terminal includes:
[0029] Call a deep learning framework according to the type of the operating system running to load the machine learning model stored in the terminal;
[0030] If it is an iOS operating system, use the first framework to load the machine learning model. If it is an Android operating system, use a second framework different from the first framework to load the machine learning model.
[0031] In a second aspect, the present application also provides an electronic device, which includes at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data retrieval and management method based on the terminal model in the first aspect.
[0032] In the above data retrieval and management method based on the terminal model, a first interface is provided to obtain the retrieval content input by the user, where the retrieval content is a user-defined retrieval object; the machine learning model stored in the terminal is loaded, and the model is used to access the terminal cache data to determine whether the retrieval object exists in the cache data. The cache data includes data cached and loaded by applications, pre-built data, and data downloaded by the user; a second interface is provided to display an optional list of target data and management buttons, where the target data is the cache data retrieved by the model and containing the retrieval object; in response to the user's operation on the button, the selected data is managed, including deleting or downloading the selected data. This method can retrieve multimodal cache data through the machine learning model, enabling the user to quickly find the required cache data and display it for the user to manage, improving the retrieval efficiency and management convenience. Description of the Drawings
[0033] Figure 1 It is an application scenario diagram of the data retrieval and management method based on the terminal model in the embodiment of the present application;
[0034] Figure 2 It is a flowchart of the data retrieval and management method based on the terminal model in the embodiment of the present application;
[0035] Figure 3 It is a flowchart of data retrieval when the retrieval object is in the form of a picture or video in the embodiment of the present application;
[0036] Figure 4 It is a flowchart of data retrieval when the retrieval object is in the form of audio in the embodiment of the present application;
[0037] Figure 5 It is a flowchart of determining the matching degree between the first feature vector and multiple second feature vectors in the embodiment of the present application;
[0038] Figure 6 It is a flowchart of determining the matching degree between the third feature vector and multiple fourth feature vectors in the embodiment of the present application;
[0039] Figure 7 This is a flowchart for loading and storing a machine learning model in the embodiments of the present application;
[0040] Figure 8 This is a structural diagram of an electronic device in the embodiments of the present application. Detailed implementation manners
[0041] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0042] As Figure 1 shown in the application environment, the application environment includes a terminal 102 and a cloud 104. The data retrieval and management method based on the terminal model provided by the embodiments of the present application is applied to the terminal 102. Among them, the terminal 102 communicates with the cloud 104 through a network. Multiple machine learning models with different parameter sizes can be stored in the cloud 104. A suitable machine learning model is downloaded from the cloud 104 based on the cache data attributes and application environment in the terminal 102. Further, the user can retrieve and manage the required cache data through the machine learning model stored in the terminal 102. Among them, the terminal 102 can be a smart phone, a tablet computer, etc. The cloud 104 can be implemented by an independent server or a server cluster composed of multiple servers. It should be noted that the data retrieval and management method based on the terminal model in the following embodiments is executed on the terminal 102 and will not be elaborated later.
[0043] In one embodiment, as Figure 2 shown, a data retrieval and management method based on a terminal model is provided. The method includes the following steps:
[0044] Step 201: Provide a first interface and obtain the retrieval content input by the user on the first interface; among them, the retrieval content includes a retrieval object defined by the user;
[0045] The user can input the content to be retrieved on the first interface. The retrieval content can include a retrieval object defined by the user. For example, the user can input a specific keyword, phrase or customize the data or specific item that the user wants to find by voice input or other means. Exemplarily, the retrieval content can be "find all documents edited last week" or "search for photos containing a specific scene", etc.
[0046] Step 202: Load the machine learning model stored in the terminal, and use the machine learning model to access the cached data of the terminal and determine whether the retrieval object exists in the cached data; wherein, the cached data is the data used when running the application program of the terminal, and the used data includes at least one of the following: the data loaded by the application cache, the pre-built data, and the data downloaded in response to the user's download instruction;
[0047] The cached data is saved in the application sandbox of the application program and is uniformly managed by the sandbox mechanism. The application sandbox is a security mechanism that can provide an independent and restricted running environment for each application program to ensure that the application program can only access the data and resources within its own sandbox.
[0048] The file formats of the cached data include but are not limited to static image or quasi-static image file formats such as JPEG, PNG, GIF, WebP, video file formats such as MP4, AVI, QuickTime, audio file formats such as MP3, ALAC, AAC, and 3D model file formats such as OBJ, FBX, usd / usdz.
[0049] It should be noted that the machine learning model can be a model learned through training data and can perform operations such as prediction, classification, or feature extraction on the input data. When accessing the cached data of the terminal, the machine learning model can use the cached data as input and process and analyze it. The machine learning model identifies the features in the cached data through its internal algorithms and parameters, matches and compares them with the features of the retrieval object, so as to determine whether there is content in the cached data that matches the retrieval object. The matching process can include operations such as feature extraction, similarity calculation, classification, or regression, depending on the type of the machine learning model and the requirements of the retrieval task.
[0050] The machine learning model can retrieve cached data in multiple formats, including pictures, videos, audio, etc. For example, for cached data in picture format, the machine learning model can extract visual features in the image; for video data, the model can extract visual features and time series features in the video frames; for audio data, the model can extract features such as the frequency, pitch, and timbre of the audio signal. Different feature extraction processes enable the machine learning model to effectively retrieve and match cached data in different formats.
[0051] Step 203: Provide a second interface, which displays an optional list of target data and buttons for managing at least some of the target data selected by the user from the optional list. The target data is the cached data containing the retrieval object retrieved by the machine learning model;
[0052] The second interface can be an interface containing operable elements. Among them, the optional list of target data can refer to a list displayed in the second interface, which can present the target data to the user in a selectable manner, such as being displayed in list form or presented as thumbnails. The button for managing at least part of the target data selected by the management user can refer to the interactive button provided by the second interface. Through these buttons, the user can perform further operations on one or more selected target data, and these buttons can provide the user with an intuitive way to manage and process the retrieved target data.
[0053] Step 204: In response to the user's operation on the button, manage at least part of the selected target data; among them, managing at least part of the selected target data includes one of the following: deleting at least part of the selected target data, downloading at least part of the selected target data.
[0054] The terminal can manage at least part of the target data selected by the user based on the user's button operation. It should be noted that the user's button operation can include a deletion operation and a download operation. Among them, the deletion operation can mean that after the user selects at least part of the target data and clicks the "Delete" button, the terminal can remove the selected at least part of the target data from the cached data to free up storage space or clean up unnecessary files. The download operation can mean that after the user selects at least part of the target data and clicks the "Download" button, the terminal can save the selected at least part of the target data to a location specified by the user for subsequent use or archiving by the user. The terminal can also perform a sharing operation on at least part of the selected data, that is, if the user wants to share the selected target data, they can click the "Share" button, which facilitates the user to transfer the target data to others or publish it on a network platform for communication and display.
[0055] In this embodiment, the terminal provides a first interface to obtain the retrieval content input by the user, where the retrieval content is a retrieval object defined by the user; loads the machine learning model stored in the terminal, uses the model to access the cached data of the terminal, and determines whether there is a retrieval object in the cached data. The cached data includes data cached and loaded by applications, pre-built data, and data downloaded by the user; provides a second interface to display an optional list of target data and management buttons, where the target data is the cached data containing the retrieval object retrieved by the model; in response to the user's operation on the button, manages the selected data, including deleting or downloading the selected data. This method can retrieve multimodal cached data through a machine learning model, enabling the user to quickly find the required cached data and display it for the user to manage, improving the retrieval efficiency and management convenience.
[0056] In one embodiment, the machine learning model includes an identification model and a matching model. Using the machine learning model to access the cached data of the terminal and determine whether there is a retrieval object in the cached data includes: using the identification model to identify the retrieval content and the cached data, and using the matching model to match the identified results to determine whether there is a retrieval object in the cached data.
[0057] Exemplarily, taking the cached data in the form of retrieved pictures as an example. For the retrieval content, the identification model can extract semantic features such as keywords and phrases; for the cached data in the form of pictures, the identification model can extract visual features such as colors, textures, and shapes. Further, the matching model is used to match the identified features. By calculating the similarity or correlation between the features, it is determined whether there is content in the cached data that matches the retrieval object. For example, if the retrieval content input by the user is "find a picture with red flowers", the identification model will extract the semantic features of "red flowers", and the matching model will search for pictures with similar visual features (such as red areas and flower shapes) in the cached data of pictures, and then determine whether there is a picture containing the detection object.
[0058] In this embodiment, the method can efficiently find the retrieval object required by the user in the cached data of the terminal through the identification model and the matching model in the machine learning model, improving the retrieval efficiency.
[0059] In one embodiment, the identification model includes a first identification model for identifying the retrieval content and a second identification model for identifying the cached data; wherein, if the format of the retrieval object is one or more of pictures, videos, and audios, the first identification model and the matching model are the same model.
[0060] When the retrieval object is in a specific format such as pictures, videos, and audios, the first identification model and the matching model in the identification model are the same model. The cached data in these formats usually has complex features and requires a powerful model to handle both the identification and matching tasks simultaneously. Take the clip model as the first identification model and the matching model. The clip model is a multi-modal model that can handle both text and visual data simultaneously. The clip model can be used as the first identification model to encode the retrieval content input by the user and extract semantic features. The clip model can also be used as the matching model to match these semantic features with the features of pictures, videos, or audios in the cached data, so as to determine whether there is content in the cached data that matches the retrieval object.
[0061] It should be noted that the main features of the CLIP model include multi-modal learning, large-scale training, and generality. It can simultaneously process image, video, and text data, map them into the same vector space, and achieve cross-modal matching and retrieval. The CLIP model can be applied to various tasks such as image classification, image retrieval, text retrieval, etc., and has strong generality and flexibility. Its working principle is to use an image encoder and a text encoder to extract features from images and texts respectively, generate image feature vectors and text feature vectors, map these feature vectors into the same vector space, and achieve matching and retrieval by calculating similarity.
[0062] In this embodiment, the method can efficiently process the retrieval task of cached data, simplify the system architecture by sharing the model structure, ensure that the recognition and matching process of cached data can be accurately carried out, and thus improve the overall retrieval efficiency and accuracy.
[0063] In one embodiment, as Figure 3 shown, the second recognition model includes an audio model for audio recognition and / or an image model for image recognition;
[0064] If the format of the retrieval object is a picture and / or a video, use the recognition model to recognize the retrieval content and the cached data, and use the matching model to match the recognized results to determine whether the retrieval object exists in the cached data, including the following steps:
[0065] Step 301: Use the text encoder of the first recognition model to encode the retrieval content to extract semantic features, obtain a first feature vector, and use the picture encoder of the image model to encode the pictures and videos in the cached data to extract visual features, obtain multiple second feature vectors;
[0066] When the format of the retrieval object is a picture and / or a video, the text encoder in the first recognition model can be used to encode the retrieval content input by the user, extract its semantic features, and generate a first feature vector. Use the picture encoder in the image model of the second recognition model to encode the pictures and videos in the cached data, extract the visual features of the pictures and videos, and generate multiple second feature vectors. The second feature vectors can represent the key visual information of the pictures and videos.
[0067] Step 302: Use the matching model to compare the first feature vector with multiple second feature vectors respectively, determine the matching degree of the first feature vector with multiple second feature vectors respectively, and determine whether the pictures and videos in the cached data contain the retrieval object according to the matching degree.
[0068] The matching model can adopt a specific algorithm to measure the matching degree between the first feature vector and each second feature vector. The higher the matching degree, the more matching the picture or video in the cached data is with the retrieved content. Further, based on the calculation result of the matching degree, it can be determined whether the pictures and videos in the cached data contain the object to be retrieved by the user.
[0069] Exemplarily, if the matching degree of a certain second feature vector and the first feature vector is higher than a preset threshold, it can be considered that the cached data corresponding to the second feature vector contains the retrieved object.
[0070] It should be noted that the first recognition model and the matching model can be chip models. When the format of the retrieved object is a picture and / or a video, the image model in the second recognition model can be the yolov8 model or the TensorFlowLite model.
[0071] In one embodiment, as Figure 4 shown, if the format of the retrieved object is audio, use the recognition model to recognize the retrieved content and the cached data, and use the matching model to match the recognized results to determine whether there is a retrieved object in the cached data, including the following steps:
[0072] Step 401: Use the first recognition model to encode the retrieved content to extract semantic features, obtaining a third feature vector, and use the audio model to encode the voice files in the cached data to extract audio features, obtaining multiple fourth feature vectors;
[0073] When the format of the retrieved object is audio, use the first recognition model to encode the retrieved content input by the user to extract its semantic features, obtaining a third feature vector. Use the audio model of the second recognition model to encode the voice files in the cached data to extract the audio features of the voice files, generating multiple fourth feature vectors, and the fourth feature vectors can represent the key audio information of the voice files.
[0074] Step 402: Use the matching model to compare the third feature vector with multiple fourth feature vectors respectively, determine the matching degrees of the third feature vector with multiple fourth feature vectors respectively, and determine whether the voice files in the cached data contain the retrieved object according to the magnitudes of the matching degrees.
[0075] The matching model can adopt a specific algorithm to measure the matching degree between the third feature vector and each fourth feature vector. The higher the matching degree, the more matching the voice file in the cached data is with the retrieved content. Further, based on the calculation result of the matching degree, it can be determined whether the voice file in the cached data contains the object to be retrieved by the user.
[0076] Exemplarily, if the matching degree of a certain fourth feature vector and the third feature vector is higher than a preset threshold, it can be considered that the cached data corresponding to the fourth feature vector contains the retrieval object.
[0077] It should be noted that the first recognition model and the matching model can be chip models. When the format of the retrieval object is audio, the audio model in the second recognition model can be the whisper model.
[0078] In this embodiment, this method can effectively handle the retrieval task of cached data, improve the accuracy and efficiency of retrieval, ensure that users can quickly and accurately find the required picture, video or audio file, and greatly improve the user experience and the convenience of data management.
[0079] In one embodiment, as Figure 5 shown, using the matching model to compare the first feature vector with multiple second feature vectors respectively to determine the matching degrees of the first feature vector with multiple second feature vectors respectively, includes: using the matching model, for any second feature vector, based on the second feature vector and the first feature vector, determining a matching degree matrix; wherein, the first feature vector includes T i text features, and any second feature vector includes I j image features, the rows of the matching degree matrix correspond to different text features T i , the columns of the matching degree matrix correspond to different image features I j , and each element T i I j in the matching degree matrix is the matching degree between the corresponding image feature and the corresponding text feature; using the matching model to determine the matching degrees of the first feature vector with each second feature vector respectively based on each matching degree matrix.
[0080] The first recognition model can encode the retrieval content input by the user, extract its text features, and obtain the first feature vector. Assume that the first feature vector contains text features. Using the picture encoder in the image model of the second recognition model to encode the pictures and videos in the cached data, extract the visual features of the pictures and videos, and generate multiple second feature vectors. Assume that any second feature vector contains I1, I2,..., I j image features. The matching model can calculate the matching degree between each text feature T i and each image feature I j , and form a T i ×I j matching degree matrix.
[0081] Exemplarily, assume that the first feature vector contains three text features T1, T2, and T3. Any second feature vector contains two image features I1 and I2. The matching model can calculate the matching degree between each text feature and each image feature to form a 3×2 matching degree matrix:
[0082]
[0083] Each element T in the matching degree matrix i I j is the matching degree between the corresponding image feature and the corresponding text feature. For example, T2I1 can represent the matching degree between the text feature T2 and the image feature I1.
[0084] The matching model can process each element in the matching degree matrix, such as calculating the maximum value, average value, or weighted average value, etc., to obtain a comprehensive matching degree score. The matching degree score can reflect the overall similarity between the first feature vector and the current second feature vector. For example, if the maximum value in the matching degree matrix is large, it indicates that at least one text feature has a high matching degree with the image feature; if the average value is high, it can indicate that the matching degree between the text features and the image features is generally high. In this way, the matching model can determine the matching degree between the first feature vector and each second feature vector, and further determine whether the pictures in the cache data contain the object retrieved by the user.
[0085] In this embodiment, this method improves the accuracy and efficiency of image data retrieval, can accurately determine whether the pictures or videos in the cache data contain the content required by the user, and optimizes the accuracy of data management.
[0086] In one embodiment, as Figure 6 shown, the matching model is used to compare the third feature vector with multiple fourth feature vectors respectively to determine the matching degrees of the third feature vector with multiple fourth feature vectors respectively, including: using the matching model, for any fourth feature vector, based on the fourth feature vector and the third feature vector, determine the matching degree matrix; wherein, the third feature vector includes T i text features, and any fourth feature vector includes I j audio features. The rows of the matching degree matrix correspond to different text features T i , and the columns of the matching degree matrix correspond to different audio features I j . Each element T in the matching degree matrix i I j is the matching degree between the corresponding audio feature and the corresponding text feature; using the matching model, based on each matching degree matrix, determine the matching degrees of the third feature vector with each fourth feature vector respectively.
[0087] The first recognition model can encode the retrieval content input by the user, extract its text features, and obtain a third feature vector. Assume that the third feature vector contains T1, T2, … T i audio features. Use the audio encoder in the audio model of the second recognition model to encode the audio files in the cached data, extract audio features, and generate multiple fourth feature vectors. Assume that any fourth feature vector contains I1, I2
[0088] …, I j audio features. The matching model can calculate the matching degree between each text feature T i and each audio feature I j , and form a matching degree matrix of T i ×I j . Each element in the matrix T i I j represents the matching degree between the corresponding text feature T i and the audio feature I j .
[0089] Exemplarily, assume that the third feature vector contains 3 text features (such as "scenery", "lake", "mountain top"), and the fourth feature vector contains 2 audio features (such as "bird song", "sound of water flow"). The matching model will construct a 3×2 matching degree matrix:
[0090]
[0091] where each element represents the matching degree between a text feature and an audio feature. Among them, the element T1I1 in the matrix can represent the matching degree between the text feature "scenery" and the audio feature "bird song", and T1I2 can represent the matching degree between the text feature "scenery" and the audio feature "sound of water flow".
[0092] The matching model can calculate all the elements in each matching degree matrix, such as calculating the average value or maximum value of all the matching degrees in the matrix, so as to quantify the similarity or correlation between the third feature vector and a specific fourth feature vector. This process can evaluate the correspondence between the text feature and the audio feature, and then judge whether the audio file in the cached data contains the content retrieved by the user. For example, if the average matching degree of a certain matching degree matrix is relatively high, it can indicate that the audio feature corresponding to the matching degree matrix is overall more matched with the text feature, and then judge that the audio in the cached data contains the object retrieved by the user.
[0093] In this embodiment, this method improves the accuracy and efficiency of audio data retrieval, can accurately judge whether the voice file in the cached data contains the content required by the user, and optimizes the accuracy of data management.
[0094] In one embodiment, the data retrieval and management method further includes: obtaining the attributes of the cached data and / or the attributes of the application environment, and obtaining a machine learning model from the cloud according to the attributes of the cached data and / or the attributes of the application environment; wherein, multiple machine learning models with different parameter amounts are stored in the cloud.
[0095] Specifically, the terminal can obtain detailed information about the cached data, such as attributes such as the type, size, format, and storage time of the cached data. The terminal can also obtain the attributes of the application environment, such as the hardware configuration of the terminal device, the network condition, and the usage habits and preferences of the user.
[0096] Based on the attributes of the cached data and / or the attributes of the application environment, the terminal can select the most suitable machine learning model for the current cached data and application environment from multiple machine learning models with different parameter amounts stored in the cloud for downloading and use. Among them, the formats of the machine learning models stored in the cloud can include TFLite and pt, etc., and the parameter amounts can be 0.5B, 1B, 3B, etc.
[0097] It should be noted that the machine learning model is updated according to new datasets every quarter to improve the performance and accuracy of the machine learning model. During the operation of the machine learning model, the terminal will regularly check the version information of the machine learning model in the cloud. If it is found that a new version of the machine learning model is released, it will automatically download and update to the latest version to ensure that the machine learning model on the terminal always remains in the latest state, thereby improving the efficiency and accuracy of data retrieval and management.
[0098] In this embodiment, the method not only improves the efficiency of data retrieval and management, but also optimizes the user experience, and can flexibly select and use machine learning models according to actual needs, so as to better meet the personalized needs of users.
[0099] In one embodiment, the data retrieval and management method further includes: providing a third interface, obtaining the set conditions and management operations for the cached data input by the user on the third interface, wherein when the cached data reaches the set conditions, the management operation for the cached data is executed.
[0100] The user can set some specific conditions on the third interface, such as the size, quantity, access times, storage time, etc. of the cached data. When the cached data reaches these set conditions, the terminal can automatically execute the management operations specified by the user, such as automatically clearing the cache, moving data to other storage locations, backing up data, etc.
[0101] In this embodiment, the method enables users to manage cache data more flexibly and efficiently, ensuring the rational use of the device's storage space. It can also automatically process cache data according to users' usage habits and needs, enhancing the user experience.
[0102] In one embodiment, the set conditions at least include one or more of the cache data exceeding a set quantity, the generation time exceeding a timeout period, the access time exceeding a set time, and the number of accesses exceeding a set number; the management operations at least include deletion and export to a preset storage location.
[0103] Exemplarily, when setting the condition that the cache data exceeds a set quantity, when the amount of cache data generated during the operation of an application reaches 1000 items set by the user, the terminal will trigger a management operation; when setting the condition that the generation time exceeds a timeout period, for example, if the generation time of the cache data exceeds 30 days from the current time, the terminal will perform corresponding processing on it; when setting the condition that the access time exceeds a set time, for example, if the cache data has not been accessed within a specified time period, such as within one month, the terminal will perform a management operation on it; when setting the condition that the number of accesses exceeds a set number, it can be that when the number of times a certain cache data is accessed by the user accumulates to a certain value, such as more than 5 times, the terminal can automatically execute relevant management operations. Among them, the management operation can be to delete the cache data or export it to a preset storage location.
[0104] In this embodiment, by setting conditions, it can be ensured that cache data is processed in a timely manner when meeting specific conditions, thereby optimizing the use of storage space and improving the efficiency and flexibility of data management.
[0105] In one embodiment, as Figure 7 shown, loading a machine learning model stored in a terminal includes the following steps:
[0106] Step 701: Call a deep learning framework according to the type of operating system running to load the machine learning model stored in the terminal;
[0107] During the operation of the terminal, according to the operating system it carries, it can select an adapted deep learning framework to load the locally stored machine learning model. This process can ensure that the machine learning model can be correctly loaded and run in different operating system environments, thereby realizing the effective retrieval and management of the cache data of the terminal. Different operating systems have different characteristics and requirements, so corresponding deep learning frameworks are needed to ensure the compatibility and performance of the model.
[0108] Step 702: If it is an iOS operating system, use the first framework to load the machine learning model; if it is an Android operating system, use a second framework different from the first framework to load the machine learning model.
[0109] When the terminal is an iOS operating system, the first framework optimized for iOS can be used to load the machine learning model. The first framework can be the CoreML framework. When the terminal is an Android operating system, the second framework used is different from the first framework, and the second framework can be the TensorFlowLite framework.
[0110] In this embodiment, the method calls the corresponding deep learning framework according to the operating system type to load the machine learning model, ensuring the efficient operation and compatibility of the model under different operating systems. This targeted framework selection and model loading method can give full play to the advantages of each operating system and improve the operation efficiency and performance of the model.
[0111] Based on the same concept, as Figure 8 shown, there is provided an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a data retrieval and management method based on a terminal model.
[0112] The electronic device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the electronic device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. The computer program, when executed by the processor, implements a data retrieval and management method based on a terminal model. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.
[0113] Those skilled in the art can understand that Figure 8 the structure shown in
[0114] It should be understood that those of ordinary skill in the art can make improvements or transformations based on the above description, and all such improvements and transformations shall fall within the protection scope of the appended claims of this application.
Claims
1. A data retrieval and management method based on a terminal model, characterized in that: The method is applied to a terminal, and the method includes: Providing a first interface, obtaining search content input by a user in the first interface; wherein the search content includes a search object defined by the user; Loading a machine learning model stored in the terminal, using the machine learning model to access cached data of the terminal, and determining whether the retrieval object exists in the cached data; wherein the cached data is data used when running an application of the terminal, and the data used includes at least one of the following: data loaded by the application cache, pre-built-in data, and data downloaded in response to a download instruction from a user; Provide a second interface, wherein the second interface displays a selectable list of target data and a button for managing at least part of the target data selected by the user from the selectable list, wherein the target data is cache data containing the retrieval object retrieved by the machine learning model; In response to the user's operation on the button, managing at least part of the selected target data; wherein managing at least part of the selected target data includes one of the following: deleting at least part of the selected target data, downloading at least part of the selected target data.
2. The data retrieval and management method based on the terminal model according to claim 1 is characterized in that: The machine learning model includes a recognition model and a matching model, and using the machine learning model to access cache data of the terminal and determine whether the retrieval object exists in the cache data includes: The search content and the cache data are identified using the recognition model, and the identified results are matched using the matching model to determine whether the search object exists in the cache data.
3. The data retrieval and management method based on terminal model according to claim 2 is characterized in that: The recognition model includes a first recognition model for recognizing the retrieved content and a second recognition model for recognizing the cached data; Among them, if the format of the retrieval object is one or more of a picture, a video, and an audio, the first recognition model and the matching model are the same model.
4. The data retrieval and management method based on the terminal model according to claim 3 is characterized in that: The second recognition model includes an audio model for audio recognition and / or an image model for image recognition; If the format of the search object is a picture and / or a video, using the recognition model to identify the search content and the cached data, using the matching model to match the recognized results, and determining whether the search object exists in the cached data, includes: Encode the search content using a text encoder of the first recognition model to extract semantic features to obtain a first feature vector, and encode the pictures and videos of the cached data using a picture encoder of the image model to extract visual features to obtain a plurality of second feature vectors; Using the matching model, the first feature vector is compared with the plurality of second feature vectors respectively, to determine the matching degree between the first feature vector and the plurality of second feature vectors respectively, and determining whether the picture and video in the cache data contain the search object according to the matching degree; and / or, If the format of the search object is audio, using the recognition model to identify the search content and the cached data, using the matching model to match the recognition results, and determining whether the search object exists in the cached data, including: Encoding the searched content using the first recognition model to extract semantic features to obtain a third feature vector, and encoding the voice file of the cached data using the audio model to extract audio features to obtain a plurality of fourth feature vectors; The third feature vector is compared with the plurality of fourth feature vectors respectively using the matching model to determine the matching degree between the third feature vector and the plurality of fourth feature vectors respectively, and whether the voice file in the cache data contains the search object is determined based on the matching degree.
5. The data retrieval and management method based on the terminal model according to claim 4 is characterized in that: Using the matching model to compare the first feature vector with the plurality of second feature vectors respectively, and determining the matching degree between the first feature vector and the plurality of second feature vectors respectively, includes: Using the matching model, for any second eigenvector, based on the second eigenvector and the first eigenvector, a matching degree matrix is determined; wherein the first eigenvector includes T i text features, any of the second feature vectors includes I j image features, the rows of the matching matrix correspond to different text features T i , the columns of the matching matrix correspond to different image features I j , each element T in the matching matrix i I j is the matching degree between the corresponding image features and the corresponding text features; The matching model is used to determine the matching degree between the first feature vector and each of the second feature vectors based on each of the matching degree matrices.
6. The data retrieval and management method based on terminal model according to claim 1 is characterized in that: The data retrieval and management method also includes: obtaining the attributes of cached data and / or the attributes of the application environment, and obtaining the machine learning model from the cloud based on the attributes of the cached data and / or the attributes of the application environment; wherein the cloud stores multiple machine learning models with different parameter sizes.
7. The data retrieval and management method based on terminal model according to claim 1 is characterized in that: The data retrieval and management method further includes: providing a third interface, acquiring setting conditions and management operations regarding the cached data input by the user on the third interface, wherein when the cached data meets the setting conditions, the management operations for the cached data are executed.
8. The terminal model-based data retrieval and management method according to claim 7, characterized in that: The set conditions include at least one or more of the cache data exceeding a set amount, the generation time exceeding a timeout period, the access time exceeding a set time, and the number of accesses exceeding a set number; and the management operations include at least deletion and exporting to a preset storage location.
9. The data retrieval and management method based on terminal model according to claim 1, characterized in that: Loading a machine learning model stored in the terminal includes: Invoke a deep learning framework to load a machine learning model stored in the terminal according to the type of operating system being run; If it is an iOS operating system, use a first framework to load the machine learning model; if it is an Android operating system, use a second framework different from the first framework to load the machine learning model.
10. An electronic device, characterized in that: include: at least one processor; And, a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the terminal model-based data retrieval and management method as described in any one of claims 1 to 9.
Citation Information
Cited By
Heterogeneous data caching method oriented to multi-modal real-time search
CN120687494A