Content query method and apparatus, and electronic device and computer-readable medium

By introducing a search platform into mobile phones, combining keyword and semantic search methods, the problem of information overload on mobile phones is solved, achieving a unified search experience and efficient retrieval results, while ensuring data security and privacy.

WO2026152876A1PCT designated stage Publication Date: 2026-07-23GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
Filing Date
2025-11-21
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

The current mobile phones contain a lot of information, and users cannot retrieve any content through a unified search portal. Furthermore, existing search methods cannot support intelligent fuzzy search using natural language, resulting in poor search results and issues with data security and privacy protection.

Method used

A search platform is provided that combines keyword and semantic search methods to obtain the content to be queried, and searches for matching business data in the accessed target business data, thereby improving the accuracy of retrieval by combining the results.

Benefits of technology

It achieves a unified search experience across different business functions, improves the relevance and accuracy of search results, ensures data security and privacy, and reduces search latency caused by network dependence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025136755_23072026_PF_FP_ABST
    Figure CN2025136755_23072026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a content query method and apparatus, and an electronic device and a computer-readable medium. The method comprises: acquiring content to be queried; on the basis of a keyword search approach and a semantic search approach, respectively searching, from among service data of at least one target service accessing a data middle platform, for service data matching said content, so as to obtain a first search result corresponding to the keyword search approach and a second search result corresponding to the semantic search approach; fusing the first search result with the second search result, so as to obtain a fused result; and on the basis of the fused result, obtaining a target query result corresponding to said content. Therefore, a data middle platform enables access operations of different services, such that different services can use the data middle platform to implement search functions, and each accessing service can have the same search experience; in addition, a keyword search approach and a semantic search approach are combined to obtain a query result of content to be queried, such that the search accuracy of the query result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Content retrieval methods, devices, electronic equipment and computer-readable media

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202510061054.7, filed on January 14, 2025, entitled "Content Search Method, Apparatus, Electronic Device and Computer-Readable Medium", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of mobile terminal technology, and more specifically, to a content retrieval method, apparatus, electronic device, and computer-readable medium. Background Technology

[0004] With the development of software and hardware technologies, mobile phones are becoming increasingly feature-rich, allowing users to enjoy a wide range of experiences solely through their devices. However, from another perspective, the ever-increasing number of applications and functions has also led to information overload on mobile phones. When users need a specific function or data, the current method usually requires them to enter a specific application and search for precise keywords to find the desired content, which is not a very efficient search method. Summary of the Invention

[0005] This application proposes a content retrieval method, apparatus, electronic device, and computer-readable medium to improve upon the aforementioned deficiencies.

[0006] In a first aspect, this application provides a content query method applied to a search platform installed in an electronic device. The method includes: acquiring content to be queried; searching for business data matching the content to be queried in business data of at least one target business accessed by the search platform, based on keyword search and semantic search respectively, to obtain a first search result corresponding to the keyword search and a second search result corresponding to the semantic search; merging the first search result and the second search result to obtain a fusion result; and obtaining a target query result corresponding to the content to be queried based on the fusion result.

[0007] Secondly, this application also provides a content query device applied to a search platform installed in an electronic device. The device includes: an acquisition unit, a search unit, a fusion unit, and a processing unit. The acquisition unit is used to acquire the content to be queried; the search unit is used to search for business data matching the content to be queried in business data of at least one target business accessed by the search platform, based on keyword search and semantic search respectively, to obtain a first search result corresponding to the keyword search and a second search result corresponding to the semantic search; the fusion unit is used to fuse the first search result and the second search result to obtain a fusion result; and the processing unit is used to obtain a target query result corresponding to the content to be queried based on the fusion result.

[0008] Thirdly, this application also provides an electronic device, comprising: one or more processors; a memory; and one or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more application programs are configured to perform the methods described above.

[0009] Fourthly, this application also provides a computer-readable medium storing processor-executable program code that, when executed by the processor, causes the processor to perform the above-described method. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 shows a flowchart of a content query method provided in an embodiment of this application;

[0012] Figure 2 shows a flowchart of a content query method provided in another embodiment of this application;

[0013] Figure 3 shows a schematic diagram of the data index creation framework provided in one embodiment of this application;

[0014] Figure 4 shows a flowchart of a content query method provided in another embodiment of this application;

[0015] Figure 5 illustrates a schematic diagram of the fusion process of multiple search methods provided in an embodiment of this application;

[0016] Figure 6 shows a flowchart of a content query method provided in another embodiment of this application;

[0017] Figure 7 shows the overall architecture diagram of the content query method provided in one embodiment of this application;

[0018] Figure 8 shows a schematic diagram of the application of the content query method provided in an embodiment of this application;

[0019] Figure 9 shows a block diagram of a content query device provided in an embodiment of this application;

[0020] Figure 10 shows a structural block diagram of the electronic device provided in an embodiment of this application;

[0021] Figure 11 illustrates a storage unit according to an embodiment of the present application for storing or carrying program code implementing the method according to an embodiment of the present application. Detailed Implementation

[0022] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. The components of the embodiments of the present application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without inventive effort are within the scope of protection of the present application.

[0023] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0024] With the development of software and hardware technologies, mobile phones are becoming increasingly feature-rich, allowing users to enjoy a wide range of experiences solely through their devices. However, from another perspective, the ever-increasing number of applications and functions has also led to information overload on mobile phones. When a user needs a specific function or data, the current method usually requires the user to enter a specific application and search for precise keywords to find that particular function or data. Furthermore, when the user is unsure which application it originated from or what the specific keywords are, they are unable to find the accurate target through searching.

[0025] In short, current mobile phone data and functions are isolated by application, preventing users from accessing any desired content on their device through a unified search portal. Furthermore, on-device local searches typically only support precise keyword-based searches or limited fuzzy searches using synonyms or near-synonyms; when users cannot recall the exact keywords, they cannot find accurate content. Generally, on-device local searches do not support intelligent fuzzy searches based on natural language.

[0026] The inventors discovered the following drawbacks in the current search function during their research:

[0027] 1) Data access-based methods require each application to access data from other applications, which is time-consuming and labor-intensive to develop and maintain; it also fails to provide a consistent search experience to other applications; and it cannot effectively guarantee data security.

[0028] 2) Based on the access method of the search interface, it is limited by the search capabilities of each application itself, and it is impossible to obtain a consistent smart search experience; it is also impossible to effectively filter, summarize and integrate the search results of each application.

[0029] 3) Keyword-based edge search methods typically only support precise keyword searches and limited generalized searches using synonyms and near-synonyms; they do not support natural language search requests, and the retrieval results are usually poor.

[0030] 4) Cloud-based natural language search can usually support fuzzy intelligent search based on natural language, thus providing a better search experience; however, it requires data transmission of business data to the cloud for indexing and storage, which cannot effectively protect user data privacy; at the same time, due to network conditions, the search latency is relatively long, which affects the search experience to some extent.

[0031] Please refer to Figure 1. Figure 1 illustrates a content query method provided in an embodiment of this application, applied to a search platform installed in an electronic device. This search platform can be an application within the electronic device, providing a unified interface, such as a unified access SDK. The application installed in the electronic device can access the search platform through this unified interface. It is understood that accessing the search platform means that the search platform has access to at least a portion of the data of the application accessing the search platform. For example, this access permission can be authorized by the application developer or by the user using the application. Then, this at least a portion of the data is applied to the content query operation. Specifically, the method includes: S101 to S104.

[0032] S101: Retrieve the content to be queried.

[0033] For example, the content to be queried can be text, voice, or an image. That is, when a user enters the content to be queried, they can input text, voice, or an image; of course, they can also input text, voice, and an image simultaneously—there is no limitation on this. It can be understood that the content to be queried can be considered a query request initiated by the user or content determined based on the user's query request. The electronic device needs to recognize the content to be queried and then return the corresponding response content, i.e., the target query result.

[0034] S102: In the business data of at least one target business accessed by the search platform, business data matching the content to be queried is searched based on keyword search and semantic search respectively, to obtain a first search result corresponding to the keyword search and a second search result corresponding to the semantic search.

[0035] It is understandable that a Data Middle Platform (DMP) is a platform that integrates search services, providing index creation and search capabilities for other businesses. The search platform corresponds to access businesses, which refer to applications or services that access the search platform, such as file, settings, and note-taking applications. The target business can be at least a portion of the access businesses corresponding to the search platform. For example, it can be a business that matches the query scope corresponding to the content to be queried. This query scope can be a range set by the user when entering the content to be queried. For example, the user sets the type of data to be queried, such as setting the query type to include all documents or setting the query type to include all video-related data. Therefore, the target business determined by the search platform can be a business that matches the access business and the query scope, such as an application with document processing capabilities or an application with video processing capabilities.

[0036] It should be noted that the business data corresponding to the target service can be data stored in the local storage space of the electronic device, data generated by the user within the target service, or data sent by other terminals received by the target service within a certain period of time; there are no restrictions on this.

[0037] As one implementation method, business data from at least one target business accessed by the search platform can be named a target dataset. Based on keyword search, business data matching the query content can be found within this target dataset to obtain the first search result. Specifically, keyword search (also known as search term-based search) is a search method whose basic principle is to compare user-inputted query terms (i.e., keywords) with words contained in the data to find documents or information containing these keywords. For example, a user inputs one or more keywords, the search platform searches for data related to the query keywords in a pre-built index, then sorts the results according to relevance and returns them to the user.

[0038] Simultaneously, it is also necessary to use semantic search to find business data matching the query content within the target dataset to obtain a second search result. Semantic search is a search method based on semantic understanding. Unlike keyword search, the goal of semantic search is to understand the semantic intent behind the query and return relevant content based on semantic similarity rather than literal matching. Semantic search relies on Natural Language Processing (NLP) technology. For example, when a user inputs the query content, the search platform first uses NLP technology to perform semantic analysis on the query to understand its underlying intent and meaning. The search platform then uses a model to transform the query into one or more semantic vectors and compares them with pre-calculated document vectors in the business data. The semantic similarity of the documents (usually calculated using cosine similarity or other distance metrics) determines their relevance to the query. Based on relevance, the results that best match the user's intent are returned.

[0039] S103: Merge the first search result and the second search result to obtain the fused result.

[0040] Understandably, the first search result represents the business data that the search platform finds by using keyword search methods and that matches the content to be queried, while the second search result fusion represents the business data that the search platform finds by using semantic search methods and that matches the content to be queried.

[0041] One implementation method is to merge the first and second search results to obtain a fused result. This allows the fused result to fully leverage the advantages of both search methods, compensate for the shortcomings of a single search approach, and improve the relevance and accuracy of the search results. Specifically, keyword search can precisely match results containing the query terms in documents, making it suitable for scenarios where the query terms are clear and unambiguous. Semantic search can understand the semantics and context of the query, providing good support for synonyms, near-synonyms, and complex queries. It can also handle fuzzy user queries and understand the underlying intent.

[0042] Understandably, merging the search results from both sources can compensate for the limitations of each. The merged search results ensure that users can accurately find content containing the keywords, while also covering semantically similar content, thus improving the completeness of the results. For example, if a user searches for "smartphone," keyword search can accurately find documents containing "smartphone," while semantic search can return results related to "phone" or "mobile device," ensuring that the user does not miss any relevant information.

[0043] S104: Based on the fusion result, obtain the target query result corresponding to the content to be queried.

[0044] In one implementation, the search platform can return the fusion result as the target query result corresponding to the content to be queried to the user. Alternatively, it can further derive the target query result corresponding to the content to be queried based on the fusion result and the search result. For example, based on the fusion result, a large language model can be combined to obtain the target query result corresponding to the content to be queried. See subsequent embodiments for details.

[0045] Therefore, in this embodiment of the application, the search platform enables access operations for different services, allowing different services to use the search platform to perform search functions, and all accessed services to obtain the same search experience; in addition, combining keyword search and semantic search to obtain query results for the content to be queried can improve the accuracy of the search results.

[0046] Please refer to Figure 2, which illustrates a content query method provided by an embodiment of this application, applied to the aforementioned search platform. Specifically, the method includes: S201 to S206.

[0047] S201: Retrieve the content to be queried.

[0048] S202: Obtain the data index corresponding to the business data of at least one target business accessing the search platform, wherein the data index includes first index information and second index information.

[0049] It should be noted that the first index information includes preset keywords corresponding to the business data, and the second index information includes preset semantic features corresponding to the business data. That is, the first index information includes the identity identifiers of multiple business data and preset keywords corresponding to each business data identity identifier. These preset keywords can be keywords determined by the business data based on a keyword inverted index. Specifically, each business data corresponds to data description information, which describes the content of the business data and can be considered a summary of the business data. By determining the preset keywords corresponding to the data description information of each business data based on the keyword inverted index, the data description information corresponding to the preset keywords can be determined. This data description information, in turn, corresponds to the identity identifier of the business data, thus allowing the determination of the preset keywords corresponding to each business data. Similarly, for the second index information, preset semantic features corresponding to each business data can also be established through semantic analysis of the data description information of the business data.

[0050] S203: Based on the keyword search method, search for business data that matches the content to be queried in the first index information to obtain the first search result.

[0051] Specifically, the query content is analyzed, keywords are extracted, and then the keywords of the query content are matched with each preset keyword in the first index information to determine the first matching degree between the query content and each business data. The first search result is obtained based on the first matching degree corresponding to each business data.

[0052] S204: Based on the semantic search method, search for business data that matches the content to be queried in the second index information to obtain the second search result.

[0053] Similarly, the semantic features of the query content are analyzed, and the semantic features of the query content are extracted. The semantic features of the query content are matched with the preset semantic features in the second index information to determine the second matching degree between the query content and each business data. The first search result is obtained based on the second matching degree corresponding to each business data.

[0054] As one implementation method, this keyword search algorithm can be a fusion of multiple different keyword search algorithms. That is, the keyword search method includes a first keyword retrieval algorithm and a second keyword detection algorithm. Based on the keyword search method, the implementation method for finding business data matching the query content in the first index information to obtain the first search result is as follows: Based on the first keyword retrieval algorithm, find business data matching the query content in the first index information to obtain a first matching result; based on the second keyword retrieval algorithm, find business data matching the query content in the first index information to obtain a second matching result; merge the first matching result and the second matching result to obtain the first search result.

[0055] It is understandable that the first keyword retrieval algorithm and the second keyword retrieval algorithm can be two different algorithms. The first matching result includes the first matching degree between the query content determined by the first keyword retrieval algorithm and each preset keyword in the first index information. The second matching result includes the second matching degree between the query content determined by the second keyword retrieval algorithm and each preset keyword in the first index information. Therefore, the first and second matching degrees corresponding to each business data can be merged to obtain the first search result. For example, by determining the weights corresponding to the first and second keyword retrieval algorithms, and based on these weights, the first and second matching degrees of each business data are weighted and summed to obtain the first search result. The first search result is a query result that merges the search results of the first and second keyword retrieval algorithms. This first search result corresponds to a keyword matching degree, and the keyword matching degree corresponding to the business data is the fusion result of the first and second matching degrees corresponding to that business data.

[0056] In one implementation, the first keyword retrieval algorithm is an exact matching algorithm, and the second keyword retrieval algorithm is a fuzzy matching algorithm. For example, the fuzzy matching algorithm is a LIKE search, and the exact matching algorithm is a keyword search.

[0057] LIKE search is a search method based on fuzzy matching in a database. It uses the LIKE operator to query whether the data contains a specific character pattern. LIKE search typically uses wildcards for partial matching. LIKE search can be considered a type of fuzzy matching; it does not require an exact match and can use wildcards for partial matching. Keyword search (sometimes called full-text search) is a precise keyword matching search of text, typically used for retrieving more complex text content. Compared to LIKE search, keyword search focuses on quickly finding words related to the query from a large amount of text, and usually utilizes techniques such as inverted indexes to improve search efficiency.

[0058] In this embodiment of the application, in the first search result, the ranking priority of the second matching result is higher than that of the first matching result; that is, the ranking priority of "like" search is higher than that of keyword search. It is understood that when displaying keyword matching results, the ranking of the second matching result is higher than that of the first matching result. Suppose that the second matching result includes four business data entries that exceed the threshold corresponding to the fuzzy matching algorithm, namely docID1, docID2, docID3, and docID4 in sequential order, while the first matching result includes two entries, namely docID3 and docID5 in sequential order. In the first search result, the four second matching results and the two first matching results can be deduplicated. That is, after deduplication, the remaining matching data includes docID1, docID2, docID3, docID4, and docID5. Regardless of whether the first matching degree of docID5 is greater than the second matching degree of docID1, docID2, docID3, and docID4, docID5 is ranked after docID1, docID2, docID3, and docID4. In other words, the first search result is ranked as docID1, docID2, docID3, docID4, and docID5. Furthermore, in the keyword matching degree corresponding to the first search result, docID1, docID2, docID3, docID4, and docID5 decrease in that order. Specifically, by reasonably setting the weights of the first keyword retrieval algorithm and the second keyword retrieval algorithm, for example, by decreasing the weight of the first keyword retrieval algorithm and increasing the weight of the second keyword retrieval algorithm, the weight of the first keyword retrieval algorithm can be made less than the weight of the second keyword retrieval algorithm. This results in the first search result having a lower fusion keyword matching degree than the second matching result, and consequently, the ranking of the matched business data in the first matching result is lower than the ranking of the matched business data in the second matching result.

[0059] It should be noted that the fusion of the first keyword retrieval algorithm and the second keyword retrieval algorithm can be referred to in the following embodiments.

[0060] As one implementation method, when obtaining a data index corresponding to the business data of at least one target service accessing the search platform, it is first determined whether the data index has already been established. If the data index has been established, it can be used directly; if not, it can be created immediately or later. Specifically, it is determined whether a data index corresponding to the business data of at least one target service accessing the search platform has already been established; if it has been established, the data index corresponding to the business data of at least one target service accessing the search platform is obtained; if not, the data index corresponding to the business data of at least one target service accessing the search platform is established.

[0061] In this embodiment, the access service synchronizes data with the search platform so that the search platform can better query data within the access service. Therefore, when the access application synchronizes data with the search platform, it also creates a data index. This typically occurs before data retrieval or immediately after a retrieval is triggered. The process of building and maintaining the index is lengthy and involves numerous modules.

[0062] Please refer to Figure 3, which illustrates the framework for creating a data index. The timing for maintaining this data index (including creation and updates) can include: proactive triggering by access services, triggering based on search requests, triggering by file explorer, and routine maintenance.

[0063] Specifically, the active triggering by the access service refers to the access service actively maintaining its index library through the SDK. For example, the maintenance of this index library includes adding, deleting, and modifying indexes for the data index of the access service. In this embodiment, these operations can be uniformly named data index update operations. For example, services such as notes and photo albums are suitable for this triggering time. The timing of index creation in this process is entirely controlled by the access service. Additionally, triggering based on a search request means that when the access service initiates a search request, the DMP has no index data. For example, the user has cleared the DMP data in application management, or has never used the relevant function. In this case, the DMP will enter the initial index building process. Furthermore, triggering based on the file explorer means that the file explorer (Metis) detects changes in external resources in real time, such as adding a new file on the phone. The file explorer (Metis) will trigger the DMP to receive file change events and enter an immediate index maintenance task. Furthermore, routine maintenance refers to the daily tasks set by the DMP that trigger daily tasks during off-peak hours. In this case, the DMP will clean up invalid indexes and continue unfinished index tasks and index maintenance tasks. Invalid indexes refer to data corresponding to the index that is invalid data, which may be data that has already been deleted.

[0064] Therefore, there are two ways to integrate an application with a DMP: semi-managed and fully managed. Semi-managed integration means that the application itself builds and maintains the index (adding, deleting, modifying, etc.), while the DMP handles routine index maintenance (cleaning up invalid data, updating keywords and vector indexes, etc.). Fully managed integration entrusts all indexing work to the DMP, which handles index creation and maintenance. Fully managed integration requires the application to provide a data acquisition interface so that the DMP can read relevant data. In essence, when integrating with a search platform, the application can specify the scope of its search, whether it needs to search only its own application data or the entire global database. The search platform will implement access control and data isolation based on the integration method.

[0065] Applications that access the DMP customize the index database table structure, namely the metainfo table, based on their own data characteristics. The metainfo table is a database table used to store and manage metadata information about the data. It typically contains attributes, characteristics, and other auxiliary information describing the data (such as title, author, tags, creation time, file type, etc.), rather than directly storing the data itself. This metadata enables more efficient data retrieval, management, and analysis. In other words, when storing data in the DMP, applications typically determine the required metadata fields based on their own needs and data characteristics.

[0066] In addition, the access application defines the columns that need to be used for keyword and vector searches based on the metainfo table structure. For example, depending on the needs of the access application, the DMP system allows indexes to be configured for certain columns in the table. These columns may include columns corresponding to fields such as title, content, author, and tags, which can be indexed for keyword searches. Each of the fields title, content, author, and tags corresponds to one column.

[0067] Furthermore, for columns with excessively long text, segmentation parameters need to be configured. In other words, for very long text columns (such as `content`), the application can configure segmentation parameters to divide the text into multiple smaller chunks. These text chunks will be stored in the `chunkinfo` table, which is associated with the `metainfo` table, for later retrieval.

[0068] Understandably, the DMP first synchronizes the raw data to the `metainfo` table. Then, based on the configuration, it segments the columns that need to be segmented into text chunks, cutting the text into chunks of the required length and storing them in the `chunkinfo` table. Next, it creates keyword inverted indexes for columns requiring keyword searches and vector indexes for columns requiring vector searches, all according to the configuration. Here, the aforementioned "text" refers to the data in certain columns of the `metainfo` table. These columns may contain long text fields (such as descriptions, notes, comments, content, etc.), i.e., the descriptive information of certain columns.

[0069] Specifically, the DMP first synchronizes the raw data provided by the application (such as logs, files, event data, etc.) to the `metainfo` table. At this point, the data is stored directly in its raw format. Then, the data is converted to a format conforming to the `metainfo` table structure. The `chunkinfo` table is used to store the segmented text blocks. For columns marked as needing segmentation, the DMP divides its text data into blocks and stores each text block in the `chunkinfo` table. Each text block may contain a portion of the original text, and there is a certain order identifier between blocks, with each text block corresponding to a position in the original text. This segmentation improves retrieval efficiency and avoids processing extremely long texts.

[0070] In the `metainfo` table, DMP creates inverted indexes for columns requiring keyword searching (such as `event_type`, `device_type`, etc.) based on configuration requirements. This allows the system to quickly find data records matching the query terms during retrieval. Specifically, a keyword inverted index records all documents or entries containing a given keyword. For columns requiring vector retrieval, DMP creates dedicated vector indexes. Vector indexes are used to efficiently find highly similar vectors and typically employ specific algorithms (such as Hierarchical Navigable Small World (HNSW) or Inverted File (IVF)) to optimize retrieval efficiency.

[0071] It should be noted that for indexes on columns that need to be partitioned, the index will be created on the partitioned data rather than the original data. Similarly, retrieval will be mapped back to the original metainfo data item after the partition is retrieved. DMP's daily tasks will clean up invalid index data and continue unfinished data indexing and index maintenance tasks.

[0072] In other words, the index is applied to the chunked data. When original text data (such as a long text field in a column) is chunked, it is divided into multiple smaller chunks, each of which may be stored in a different record. Although the original text field may be long, the chunked data is actually composed of multiple shorter text blocks. In this case, to improve search or query efficiency, the index is built on the chunked data, rather than directly on the original long text field. It's understandable that chunked text data is generally more suitable for indexing than the original long text because each block is shorter, allowing for more efficient full-text indexing, keyword searching, and other operations. For example, suppose there is a long text field named `event_description`, which has been configured to be chunked into multiple smaller chunks and stored in the `chunkinfo` table. For each smaller chunk (let's say named `chunk_text`), and assuming `event_description` is divided into three chunks and stored in the `chunkinfo` table, an index can be created on the `chunk_text` field of the `chunkinfo` table, allowing for faster searches of each chunk during queries. Without segmentation, a full-text index might be directly created on the original `event_description` field. This would require the search engine to process longer text during queries, potentially impacting performance. For example, if a text is segmented into segment 1 and segment 2, a local search would first retrieve segment 1, then map it to the original text and the location of the corresponding specific text occurrences.

[0073] Furthermore, when creating and maintaining data indexes, it's necessary to consider whether the current state of the electronic device is suitable for the creation or maintenance of the data index. In other words, the creation or maintenance of the data index requires determining whether the electronic device meets the corresponding conditions. Specifically, the DMP creates data indexes for local note-taking settings in batches, and during each batch of data processing, it considers factors such as phone battery level, storage, and screen status to reduce power consumption and the impact of phone overheating and lag.

[0074] Understandably, considering that the workload is relatively large when creating a data index for the first time, and it is usually triggered by the user, different conditions can be set for the first creation and non-first creation to determine whether the objective conditions for performing the data index creation operation are met.

[0075] Regarding the data index creation process, upon detecting a data index creation request, it is determined whether this is the first creation request. If it is the first creation request, the data index creation operation is executed if the electronic device meets the first condition. If it is not the first creation request, the data index creation operation is executed, and if, during the execution of the data index creation operation, it is detected that the electronic device meets the second condition, the current data index creation operation is interrupted. In other words, for data index creation operations that are not the first, the data index creation operation can be executed first. During execution, it is determined whether the electronic device meets the second condition. If the second condition is met, the current data index creation operation is aborted, and then execution continues only after the next task trigger. It is understood that the data index creation request can be triggered by the user or the application; please refer to the aforementioned explanation of triggering timing for details.

[0076] As one implementation method, when setting the first condition for the initial creation of a data index, it is also necessary to consider whether the current time is during the daytime or nighttime period. This is because the state requirements of the electronic device for the data indexing operation are different during the daytime and nighttime periods. The setting of the first condition not only needs to take into account the power consumption of the electronic device, but also needs to ensure that the data index creation operation can be successfully completed.

[0077] Specifically, if this creation request is the first creation, the target time period corresponding to the current time is determined. If the target time period is within the daytime period, the first condition is determined to include at least one of the following: the remaining power of the electronic device is greater than a first power threshold, the casing temperature of the electronic device is less than a first temperature threshold, the CPU load of the electronic device is less than a first load threshold, and the remaining storage space of the electronic device is greater than a first space threshold.

[0078] Both the daytime and nighttime time periods can be set based on actual usage needs. For example, they can be set based on the user's lifestyle habits. These habits can be manually set by the user or, with the user's authorization, set by analyzing the user's electronic device usage data. For example, the time period when the user uses their mobile phone most frequently during the day is the daytime time period, without limitation. For instance, the first battery threshold, first temperature threshold, first load threshold, and first space threshold can also be set based on actual usage needs. For example, the first battery threshold could be 80% (i.e., 80% of the maximum battery capacity), the first temperature threshold could be 41 degrees Celsius, the first load threshold could be 80% (80% refers to 80% CPU utilization, while 100% CPU load means all CPU cores are fully occupied), and the first space threshold could be 500MB.

[0079] It is understood that if the time of obtaining the initial index creation request is during daytime, the corresponding first condition can be at least one of the following: the remaining battery power of the electronic device is greater than a first battery power threshold, the casing temperature of the electronic device is less than a first temperature threshold, the CPU load of the electronic device is less than a first load threshold, and the remaining storage space of the electronic device is greater than a first storage space threshold. That is, it can be any one of them or a combination of multiple ones. In the embodiments of this application, the first condition corresponding to the daytime period can include the remaining battery power of the electronic device being greater than the first battery power threshold, the casing temperature of the electronic device being less than the first temperature threshold, the CPU load of the electronic device being less than the first load threshold, and the remaining storage space of the electronic device being greater than the first storage space threshold. That is, the electronic device is determined to meet the first condition if all of the following conditions are met: the remaining battery power of the electronic device is greater than the first battery power threshold, the casing temperature of the electronic device is less than the first temperature threshold, the CPU load of the electronic device is less than the first load threshold, and the remaining storage space of the electronic device is greater than the first storage space threshold.

[0080] Understandably, if the target time period falls within nighttime, the first condition includes at least one of the following: the electronic device is idle; the electronic device is in a screen-off state; the electronic device is charging and its remaining battery level is greater than a second battery threshold; the electronic device's casing temperature is less than a first temperature threshold; the electronic device's CPU load is less than a first load threshold; and the electronic device's remaining storage space is greater than a first storage space threshold. If the electronic device meets the first condition, the data index creation operation is executed. Similarly, the second battery threshold can also be set based on actual usage; for example, the second battery threshold could be 20%. The first condition corresponding to the nighttime period can be at least one of these conditions, or a combination of several. In this embodiment, the first condition corresponding to the nighttime period can include the electronic device being in an idle state, the electronic device being in a screen-off state, the electronic device being in a charging state with remaining battery power greater than a second battery power threshold, the electronic device having a casing temperature less than a first temperature threshold, the electronic device having a CPU load less than a first load threshold, and the electronic device having remaining storage space greater than a first storage space threshold. That is, if all of the following conditions are met during the nighttime period, the electronic device is determined to meet the first condition. The nighttime period can be 2-5 AM.

[0081] Then, after determining the first condition, and if the electronic device satisfies the first condition, the data index creation operation is performed.

[0082] As one implementation, if the detected data index creation request is not the first time it has been created, a second condition needs to be determined based on the current time period of the electronic device. Specifically, if it is not the first time it has been created, the target time period corresponding to the current time is determined; if the target time period is within the daytime period, the second condition is determined to include the electronic device's casing temperature being greater than or equal to a first temperature threshold, the temperature difference of the electronic device's casing temperature rise during the data index creation operation being greater than a second temperature threshold, or the electronic device's power consumption during the data index creation operation exceeding a second power threshold; if the target time period is within the nighttime period, the second condition is determined to include the electronic device's casing temperature being greater than or equal to the first temperature threshold, the electronic device's power consumption during the data index creation operation exceeding a second power threshold, or the electronic device being in a screen-on state; the data index creation operation is executed, and if it is detected that the electronic device meets the second condition during the execution of the data index creation operation, the current data index creation operation is interrupted.

[0083] It is understandable that the second temperature threshold and the second power threshold can also be set based on actual usage needs. For example, the second temperature threshold could be 2 degrees Celsius, and the second power threshold could be 2%. In cases other than the first creation, if the current time is during daytime, the electronic device can be determined to meet the second condition if any one of the following conditions is met: the electronic device's casing temperature is greater than or equal to the first temperature threshold; the temperature difference of the electronic device's casing temperature rise during the data index creation operation is greater than the second temperature threshold; and the electronic device's power consumption during the data index creation operation exceeds the second power threshold. Conversely, in cases other than the first creation, if the current time is during nighttime, the electronic device can be determined to meet the second condition if any one of the following conditions is met: the electronic device's casing temperature is greater than or equal to the first temperature threshold; the electronic device's power consumption during the data index creation operation exceeds the second power threshold; and the electronic device is in a screen-on state.

[0084] As another implementation, for the data index update operation, if an update request for an established data index is detected, it is determined whether the update request is the first update; if it is the first update, the data index update operation is performed if the electronic device meets a third condition; if it is not the first update, the data index update operation is performed if the electronic device meets a fourth condition.

[0085] It is understood that the data index update operation can refer to operations such as adding or modifying an existing data index. If this is the first update, the third condition is determined to include at least one of the following: the remaining battery power of the electronic device is greater than a second battery threshold, the CPU load of the electronic device is less than a first load threshold, and a search operation has been performed based on the data index within a specified time period. If the electronic device meets the third condition, the data index update operation is performed. That is, the third condition can be at least one of several conditions, such as the remaining battery power of the electronic device being greater than the second battery threshold, the CPU load of the electronic device being less than the first load threshold, and a search operation has been performed based on the data index within a specified time period, or any combination of these conditions. In this embodiment, the third condition is that the remaining battery power of the electronic device is greater than the second battery threshold, the CPU load of the electronic device is less than the first load threshold, and a search operation has been performed based on the data index within a specified time period, all of which are simultaneously met. The phrase "a search operation has been performed based on the data index within a specified time period" means that within a specified time period, the electronic device performed a query operation within the data index to be updated based on a user's query request; the purpose of this is to ensure that the data index to be updated has been used. The specified time period can be within 7 days of the current moment, 7 consecutive days before the current moment, or 7 consecutive days including the current day.

[0086] If it is not the first update, the fourth condition is determined to include at least one of the following: the time difference between the current moment and the last time the data index update operation was performed is greater than a duration threshold; the number of times the update operation was performed within the current time period corresponding to the current moment is less than a number threshold; the data to be updated is less than a specified quantity threshold; and the current moment is within a nighttime period. If the electronic device meets the fourth condition, the data index update operation is performed. The duration threshold, number threshold, and specified quantity threshold can be set based on actual usage. For example, the duration threshold is 10 seconds, the number threshold is 4, and the specified quantity threshold is 50. That is to say, for non-first updates, it is limited to triggering once every 10 seconds (file), no more than four times and no more than 50 documents / notes per minute (<20 seconds, <2 mAH), and updates are performed during off-peak hours at night.

[0087] S205: Merge the first search result and the second search result to obtain the fused result.

[0088] S206: Based on the fusion result, obtain the target query result corresponding to the content to be queried.

[0089] It should be noted that any content not described in detail in the above steps can be referred to the foregoing embodiments, and will not be repeated here.

[0090] Therefore, it can be seen that the retrieval method provided in this application embodiment is a hybrid retrieval that includes multiple different search algorithms. Specifically, the implementation method of the hybrid retrieval can be referred to in the following embodiments.

[0091] Please refer to Figure 4, which illustrates a content query method provided by an embodiment of this application, applied to the aforementioned search platform. Specifically, the method includes: S401 to S406.

[0092] S401: Retrieve the content to be queried.

[0093] S402: In the business data of at least one target business accessed by the search platform, business data matching the content to be queried is searched based on keyword search and semantic search respectively, to obtain a first search result corresponding to the keyword search and a second search result corresponding to the semantic search.

[0094] S403: Determine the first weight corresponding to the first search result and the second weight corresponding to the second search result.

[0095] S404: Based on the first weight and the second weight, reorder multiple first matching data and second matching data to obtain a mixed sorting result, wherein the mixed sorting result includes multiple matching data ordered sequentially.

[0096] S405: Take the top N matching data in the mixed sorting result as the fusion result, where N is an integer greater than 1.

[0097] As one implementation method, the method of fusing the first search result and the second search result is as follows: the first matching data and the second matching data are weighted and summed to recalculate the score of each matching data to obtain the score value of each matching data. For example, the score value of each matching data is obtained by weighted fusion operation based on the first weight, the keyword matching degree corresponding to the first matching data, the second weight, and the semantic matching degree corresponding to the second matching data; the multiple matching data are sorted based on the score value, and the top N matching data are taken as the fusion result, where N is an integer greater than 1.

[0098] It should be noted that the first matching data in the first search result refers to the successfully matched data determined based on this keyword search method. For example, based on this keyword search method, the keyword matching degree between each business data and the query content can be determined. Based on the keyword matching degree, the business data are sorted, and the first number of business data in the ranking is taken as the first matching data. Thus, the first matching data corresponds to the keyword matching degree. Similarly, the second matching data in the second search result refers to the successfully matched data determined based on this semantic search method. It can be understood that the first matching data and the second matching data can contain the same data, but they represent different names. For example, a text docID is recorded as the first matching data in the first search result and as the second matching data in the second search result.

[0099] As one implementation method, a weighted fusion operation is used to obtain the score value corresponding to each matching data based on the first weight, the keyword matching degree corresponding to the first matching data, the second weight, and the semantic matching degree corresponding to the second matching data. The identity identifier corresponding to the first matching data is determined, and this identity identifier is recorded as the identity information of the first matching data. For example, the aforementioned docID is the identity identifier. Thus, the keyword matching degree and semantic matching degree corresponding to each business data can be determined. Then, for each business data, its corresponding score value is the product of the first weight and the keyword matching degree, and the sum of the product of the second weight and the semantic matching degree. The result is the score value corresponding to the business data.

[0100] It is understandable that the first matching data in the first search result may be the filtered business data after calculating the keyword matching degree for each business data. That is, the aforementioned sorting of each business data based on the keyword matching degree, and taking the first number of business data at the top of the sort as the first matching data. Similarly, the second matching data in the second search result is also the filtered business data. Therefore, when merging the first search result and the second search result, some matching data may not have a corresponding keyword matching degree or semantic matching degree. In this case, the matching degree of the unmatched data can be set to zero.

[0101] In one implementation, the first weight and the second weight can be the same or different. If the first weight and the second weight are different, the second weight can be set to be greater than the first weight. Alternatively, the first weight and the second weight can be determined based on the first search result. Specifically, the implementation of determining the first weight corresponding to the first search result and the second weight corresponding to the second search result involves determining a reference matching degree based on the keyword matching degree of at least one of the keywords in the first search result; if the reference matching degree is less than a preset threshold, the second weight is set to be greater than the first weight; if the reference matching degree is greater than or equal to the preset threshold, the first weight and the second weight are set to be the same. The reference matching degree can be the average of the keyword matching degrees corresponding to each first matching data in the first search result, or it can be the maximum value of the keyword matching degrees corresponding to each first matching data. That is, the implementation of determining the reference matching degree based on the keyword matching degree of at least one of the keywords in the first search result can be to determine the largest keyword matching degree in the first search result as the reference matching degree.

[0102] As another implementation, multiple first and second matching data can be reordered based on a first weight and a second weight to obtain a mixed ranking result. The first and second weights affect the ranking position of their respective matching data in the mixed ranking result. For example, the reordering can be performed based on the following formula (1): rrf_score=weight*(1 / (rank+c)) (1)

[0103] In formula (1), c is the fusion constant, which can be set to 60, rank is the sequence number of the matched data in its corresponding queue, and weight is the first weight or the second weight.

[0104] In other words, the first search result includes multiple first matching data items in sequential order, and the second search result includes multiple second matching data items in sequential order.

[0105] For the first search result, the weight in formula (1) is the first weight. Each first matching data in the first search result has a corresponding sequence number and identity identifier in the first search result. For the first search result, each matching data can be traversed to calculate the rrf_score, i.e., the first score, of each matching data in the recall. Taking the first matching data in the first search result as documents as an example, each document corresponds to a docID, and each document corresponds to a sorting sequence number in the first search result. Based on the above formula (1), the rrf_score corresponding to each docID can be obtained. Similarly, for the second search result, the weight in formula (1) is the second weight. It can also obtain the rrf_score, i.e., the second score, of each matching data in the second search result. The first score and the second score of each matching data are added together, that is, the rrf_score in the two recalls is added to the corresponding docID to obtain the score value corresponding to each matching data. Then, the matching data is sorted based on the score value of each matching data to obtain the mixed ranking result. Then, the top N matching data in the mixed sorting results are taken as the fusion result, where N is an integer greater than 1.

[0106] S406: Based on the fusion result, obtain the target query result corresponding to the content to be queried.

[0107] As can be seen, the content query method provided in this application uses a hybrid retrieval method combining multiple search techniques. Specifically, after receiving a retrieval request, the DMP first performs word segmentation on the retrieval query. The resulting terms include time terms and ordinary terms. Time terms can be used for time-based conditional retrieval, while ordinary terms are mainly used for keyword search. For example, assuming that the keyword search method in this application includes two different keyword search algorithms—exact matching and fuzzy matching (e.g., LIKE search)—the exact matching algorithm here is a precise search, i.e., a keyword-based exact matching algorithm, compared to LIKE search. Therefore, this keyword search method aims to merge the search results of exact and fuzzy searches. Ordinary terms are mainly used in the exact matching algorithm and LIKE search, and are also generalized through synonyms and near-synonyms to achieve keyword-based fuzzy search.

[0108] When a front-end application requests a DMP to perform a search, the DMP will perform keyword retrieval and vector semantic retrieval from different resource content according to the specified search scope (such as settings, notes, files, etc., which can specify cross-application multi-resource retrieval or limit the retrieval to a single resource), and then perform a mixed sorting.

[0109] As one implementation method, it is assumed that the keyword search method mentioned in this application (i.e., traditional retrieval) includes a first keyword retrieval algorithm and a second keyword retrieval algorithm, wherein the first keyword retrieval algorithm is an exact matching algorithm, the second keyword retrieval algorithm is a fuzzy matching algorithm, and the voice search method is a vector search algorithm. Please refer to Figure 5, and the fusion process of these multiple search methods is shown in Figure 5.

[0110] In Figure 5, keyword search uses the exact matching algorithm described above, while "like" search uses the fuzzy matching algorithm described above. The matching degree of the "like" search results is calculated as the proportion of the matched text to the original text, i.e., the hit ratio. The hit ratio refers to the ratio of the portion of the query text (or keyword) appearing in the target text to the total length of the target text. For example, if the search term is "apple" and the target text is "I love apple pie", the matched portion is "apple". The hit ratio can be calculated by dividing the number of characters of "apple" appearing in the text by the total number of characters in the target text. Furthermore, the keyword matching degree of the exact matching algorithm, i.e., the ordinary keyword matching algorithm, is calculated as TF-IDF score * number of matched terms. The exact matching algorithm is commonly used to calculate the matching degree of keywords. Its core idea is to measure the importance of each term based on keyword frequency (TF) and inverse document frequency (IDF), and to obtain the final matching degree through a weighted score of these importance values. TF (Term Frequency) refers to the frequency of a word appearing in a document. A high TF value indicates a higher frequency of occurrence in the document, usually indicating that the word is more important in the current document. IDF (Inverse Document Frequency) represents the importance of a word in the entire document set. Words with high IDF values ​​usually appear less frequently in many documents, indicating that they are relatively rare in the corpus and therefore have higher weights.

[0111] As can be seen, the entire search process consists of two parts: keyword search and semantic search.

[0112] Regarding the keyword search section:

[0113] In the resource database table, a LIKE fuzzy search is performed on the relevant keyword search column. The relevant keyword search column refers to the column in the index database of the application being queried where keyword searches are performed. For example, if the application is a settings application, and the query is for information about a specific setting, then the relevant keyword search column could be a column in the index database corresponding to the settings application, such as the setting description column. In other words, a LIKE fuzzy search is performed on the keywords corresponding to a specific search column of the application being queried within the data index.

[0114] Perform keyword searches and exact matches in the relevant keyword search column of the resource database table. The explanation of the relevant keyword search column can be found in the previous content, and will not be repeated here.

[0115] Taking a scenario where the app to be queried is a settings app, and the query content is a settings item, in the keyword exact match search method, only the original word segmentation results from the word segmentation results are used to perform keyword retrieval. Synonyms from the word segmentation results and time and location terms in the NER are not used. This is because mobile phone settings queries do not have time and location characteristics; for example, a user would not say "yesterday's XX settings".

[0116] After performing a fuzzy "like" search and a precise keyword match search, the search results for the precise keyword match search and the fuzzy "like" search are obtained. Then, filtering operations are performed on these search results to obtain the first matching result for the precise keyword match search and the second matching result for the fuzzy "like" search. The filtering operation is used to filter out non-existent settings. Since the DMP does not immediately delete related data indexes when data changes, but marks them as invalid, a cleanup task is performed uniformly during off-peak hours to avoid affecting user experience. During the search, invalid data needs to be filtered out from the search results, i.e., invalid search results are filtered out. Invalid means that the data corresponding to the search result has been deleted. Then, the first matching result for the precise keyword match search and the second matching result for the fuzzy "like" search are deduplicated and merged. The merged result is named the first search result. The merging operation of the first and second matching results can be referred to the previous content and will not be repeated here.

[0117] After merging, the first search result stores a first preset number of matching data. As mentioned above, the first search result includes multiple first matching data and the keyword matching degree corresponding to each first matching data. The number of first matching data is the first preset number. For example, the first preset number is 6. That is to say, the results of like search and keyword exact match search are merged, and the merged result retains the top 6 search results.

[0118] As shown in Figure 5, the semantic search method is vector semantic search. A vector index is built based on the relevant vector index columns of the resource database table, and vector retrieval is performed in the corresponding vector index table. The creation of the vector index can refer to the previous embodiment, and will not be repeated here. Embedding vector scores can be used to score the vector semantic search. Specifically, in the resource database table, the embedding column is used to store the embedding vector of each resource (i.e., business data). After the vector index is built, vector retrieval can be performed. During retrieval, the similarity between the query vector and the embedding vector of each resource in the database needs to be calculated. In vector retrieval, the score is usually based on the similarity between the query vector and the embedding vector in the database. This yields the similarity score, i.e., the semantic matching degree, for each business data corresponding to the vector semantic search. Then, a filtering operation is performed to filter out non-existent business data (e.g., settings items), thereby obtaining the second search result.

[0119] After obtaining the first and second search results, a fusion strategy is executed. For example, RRF fusion can be used, which uses the inverse of the ranking ordinal number to merge the results of different retrieval strategies.

[0120] Specifically, the recall strategy weights and fusion constants are initialized. Three weighting strategies are used: an adaptive strategy (applied to settings) checks the maximum matching degree in traditional search; if it's lower than a set weighting threshold, the weights for traditional search and vector search are set to [0.3, 0.7], otherwise [0.5, 0.5]. The maximum matching degree in traditional search refers to the aforementioned reference matching degree. The weighting threshold can be the aforementioned preset threshold. Therefore, the weighting settings for values ​​below and above the set weighting threshold can refer to the previous content and will not be repeated here. Specifically, when the second weight is set to be greater than the first weight, the second weight can be set to 0.7 and the first weight to 0.3. Other weighting strategies include: a fixed average strategy (fixed to [0.5, 0.5]), used for chunked text in document search; and a priority strategy (fixed weight for one path at 0.7), used for title text in document search. For example, the second weight corresponding to a vector can be set to 0.7.

[0121] The implementation method based on the first weight, the keyword matching degree corresponding to the first matching data, the second weight, and the semantic matching degree corresponding to the second matching data through weighted fusion operation can be as follows: calculate the scores of docID in the two-way recall based on the above formula (1), and sum the calculation results of the two-way recall to the corresponding docID. If the evaluation uses the reranker model for re-ranking, the results of the multi-way recall are deduplicated and directly merged, and sorted according to the score of the reranker model.

[0122] Therefore, the embodiments of this application can integrate multiple retrieval methods to obtain the final retrieval results, achieving a more intelligent and accurate retrieval effect.

[0123] Please refer to Figure 6, which illustrates a content query method provided by an embodiment of this application, applied to the aforementioned search platform. Specifically, the method includes steps S601 to S604.

[0124] S601: Retrieve the content to be queried.

[0125] S602: In the business data of at least one target business accessing the search platform, business data matching the content to be queried is searched based on keyword search and semantic search respectively, to obtain a first search result corresponding to the keyword search and a second search result corresponding to the semantic search.

[0126] S603: Merge the first search result and the second search result to obtain the fused result.

[0127] S604: Using the large language model, based on the fusion result, obtain the output result of the large language model, and based on the output result, obtain the target query result corresponding to the content to be queried.

[0128] In other words, the fusion result is used as input to the large language model, and then the output result of the large language model is obtained. Based on the output result, the target query result corresponding to the query content is obtained. As one implementation method, the fusion result and the query content can be assembled into a prompt vector, and the prompt vector is input into the large language model (e.g., an LLM model).

[0129] It is understood that the large language model can be deployed on electronic devices or in the cloud. If deployed in the cloud, the implementation of S604 can be as follows: send the fusion result to the cloud, trigger the cloud to generate prompt information based on the fusion result, input the prompt information into the large language model, and obtain the target query result corresponding to the query content based on the output result of the large language model.

[0130] As shown in Figure 7, assume that the large language model is deployed in the cloud.

[0131] As shown in Figure 7, different services can access this DMP. Combining the aforementioned embodiments and applying them to Figure 7, the execution flow of this method can be:

[0132] 1) Front-end applications initiate intelligent searches through the RAGAgent interface provided by the Converged Search Service - Search Platform (DMP, hereinafter referred to as DMP);

[0133] 2) If DMP has already established the relevant data index, proceed to 8); otherwise, proceed to 3).

[0134] 3) The user is prompted that the index has not yet been created and immediate indexing is required;

[0135] 4) If the user chooses to perform the indexing later, the user will be prompted that the indexing is not yet ready, and the process will terminate.

[0136] 5) If the user selects instant indexing, proceed to 6);

[0137] 6) The DMP reads data from the access resources through the data interface;

[0138] 7) DMP performs keyword inverted indexing and vector indexing on the data of accessed resources;

[0139] 8) The DMP performs keyword inverted index retrieval and vector retrieval based on the query string passed by the front-end application;

[0140] 9) DMP merges and sorts the results of inverted index search and vector search to generate local search results;

[0141] 10) The DMP transmits local search results to the AI ​​cloud service;

[0142] 11) The AI ​​cloud service uses the Reranker model to re-rank the results;

[0143] 12) The AI ​​cloud service enhances the prompt words based on the re-ranking results and passes them to the LLM large model;

[0144] 13) The final search answer is generated in a streaming manner using a large model;

[0145] 14) The AI ​​cloud service will stream the answers back to the DMP;

[0146] 15) The DMP passes the answer to the front-end application and associates the final answer with local resources;

[0147] 16) The front-end application presents the interface.

[0148] It should be noted that the specific implementation methods of the above steps have been described in the preceding embodiments and will not be repeated here. It can be seen that, based on local search results, a large language model in the cloud can be combined to obtain more accurate response content, which can then serve as the target query result corresponding to the content to be queried.

[0149] Combining the steps shown in Figure 5 above, the fusion results of local multi-resource retrieval can be obtained. Then, the top N1 results (i.e., the top N results in the aforementioned ranking) can be retained, and the cloud interface can be called to send the search request (including the content to be queried) and the fusion results to the cloud for further retrieval and ranking.

[0150] After receiving the search request and fused search results, the cloud first uses a rerank model to sort the fused results based on the relevance of the search request and search results, retaining the top N2 results. In other words, upon receiving a request, the cloud uses a rerank model to reorder the matching data within the incoming fused results. The reranking model is typically based on machine learning or deep learning models, and its training objective is to optimize based on the relevance of the query and search results.

[0151] The filtered results are then assembled with the search request into a prompt, and the LLM interface is called to obtain the LLM output. LLM can more intelligently analyze the degree of correlation between the natural language-based search request and the searched item, thereby obtaining more accurate search results. However, LLM usually has a long processing latency, so the LLM output will be returned to the client side as streaming data for display.

[0152] Please refer to Figure 8 for an example illustrating an embodiment of this application. Figure 8 shows the interface of the All-Search application. As shown in Figure 8(a), the interface displays recommended content 801. Users can click on recommended content 801, and the clicked recommended content becomes the current query content. The cloud side can configure a daily / 7-day limit. If this limit is exceeded, the recommended content will no longer be displayed for a specified period. In addition, the cloud side pre-stores about 50 recommended words as recommended content. If the effective time exceeds this period, the project will no longer push recommended words. As shown in Figure 8(b), the interface displays a recognition control 802. After entering the query content, clicking the recognition control 802 will trigger the query operation. As one implementation method, if the input content exceeds a specified number of characters, a large model can be preloaded before the content input is completed. As shown in Figure 8(c), the interface displays a control control 803. Clicking the control control stops the generation of the target query result. The display of the target query result is shown in Figure 8(d).

[0153] Therefore, in this embodiment, the front-end access application can simultaneously search from settings, notes, documents, and other content through this solution, and can filter, summarize, and merge the results to obtain a cross-application search experience; moreover, it integrates multiple search technologies, such as time search, conditional search, keyword inverted search, and semantic vector search, and can simultaneously perform multiple search strategies on the same data source to obtain richer search results.

[0154] Furthermore, the embodiments of this application can support customized search strategies and search scopes for services. When a service accesses the application, it simultaneously specifies the search scope and search strategy, thereby achieving differentiated search results for different data characteristics of different services, better adapting to the specific needs of each service. Additionally, the aforementioned reranker model and LLM model can be further deployed on edge devices, making the entire solution completely edge-based and independent of the network, effectively avoiding situations where intelligent search is unavailable in environments without or with weak network connectivity. Moreover, after obtaining locally integrated search results, the search results can be presented directly to the front-end application instead of being transmitted to the cloud for further refinement, resulting in lower search latency and resource consumption.

[0155] Please refer to Figure 9, which shows a structural block diagram of a content query device 900 provided in an embodiment of this application. The device may include: an acquisition unit 901, a search unit 902, a fusion unit 903, and a processing unit 904.

[0156] Unit 901 is used to retrieve the content to be queried.

[0157] Search unit 902 is used to search for business data that matches the content to be queried in at least one target business data accessed by the search platform, based on keyword search and semantic search respectively, to obtain a first search result corresponding to the keyword search and a second search result corresponding to the semantic search.

[0158] Furthermore, the search unit 902 is also used to obtain a data index corresponding to the business data of at least one target business accessing the search platform, wherein the data index includes first index information and second index information, wherein the first index information includes preset keywords corresponding to the business data, and the second index information includes preset semantic features corresponding to the business data; based on the keyword search method, business data matching the content to be queried is searched in the first index information to obtain the first search result; based on the semantic search method, business data matching the content to be queried is searched in the second index information to obtain the second search result.

[0159] Furthermore, the search unit 902 is also used to search for business data matching the content to be queried in the first index information based on the first keyword retrieval algorithm to obtain a first matching result; to search for business data matching the content to be queried in the first index information based on the second keyword retrieval algorithm to obtain a second matching result; and to merge the first matching result and the second matching result to obtain the first search result.

[0160] Furthermore, the first keyword retrieval algorithm is an exact matching algorithm, the second keyword retrieval algorithm is a fuzzy matching algorithm, and in the first search result, the ranking priority of the second matching result is higher than the ranking priority of the first matching result.

[0161] Furthermore, the search unit 902 is also used to determine whether a data index corresponding to the business data of at least one target service accessing the search platform has been established; if it has been established, the data index corresponding to the business data of at least one target service accessing the search platform is obtained; if it has not been established, a data index corresponding to the business data of at least one target service accessing the search platform is established.

[0162] Furthermore, the search unit 902 is also used to determine whether the data index creation request is the first creation when a data index creation request is detected; if it is the first creation, the data index creation operation is executed if the electronic device meets the first condition; if it is not the first creation, the data index creation operation is executed, and if the electronic device meets the second condition during the execution of the data index creation operation, the current data index creation operation is interrupted.

[0163] Furthermore, the search unit 902 is also used to determine the target time period corresponding to the current moment if it is the first time it is created;

[0164] If the target time period is during the daytime, the first condition is determined to include at least one of the following: the remaining battery power of the electronic device is greater than a first battery power threshold, the casing temperature of the electronic device is less than a first temperature threshold, the CPU load of the electronic device is less than a first load threshold, and the remaining storage space of the electronic device is greater than a first storage space threshold. If the target time period is during the nighttime, the first condition is determined to include at least one of the following: the electronic device is in an idle state, the electronic device is in a screen-off state, the electronic device is in a charging state and the remaining battery power is greater than a second battery power threshold, the casing temperature of the electronic device is less than a first temperature threshold, the CPU load of the electronic device is less than a first load threshold, and the remaining storage space of the electronic device is greater than a first storage space threshold. If the electronic device meets the first condition, the data index creation operation is performed.

[0165] Furthermore, the search unit 902 is also used to determine the target time period corresponding to the current moment if it is not the first time it is created;

[0166] If the target time period is during the daytime, the second condition is determined to include: the casing temperature of the electronic device being greater than or equal to the first temperature threshold; the temperature difference during the data index creation operation being greater than the second temperature threshold; or the power consumption of the electronic device during the data index creation operation exceeding the second power threshold. If the target time period is during the nighttime, the second condition is determined to include: the casing temperature of the electronic device being greater than or equal to the first temperature threshold; the power consumption of the electronic device during the data index creation operation exceeding the second power threshold; or the electronic device being in a screen-on state. The data index creation operation is executed, and if it is detected that the electronic device meets the second condition during the execution of the data index creation operation, the current data index creation operation is interrupted.

[0167] Furthermore, the search unit 902 is also configured to determine whether the update request is the first update when an update request for an established data index is detected; if it is the first update, the data index update operation is performed if the electronic device meets the third condition; if it is not the first update, the data index update operation is performed if the electronic device meets the fourth condition.

[0168] Furthermore, the search unit 902 is also used to determine, if it is the first update, at least one of the following conditions: the remaining power of the electronic device is greater than a second power threshold, the CPU load of the electronic device is less than a first load threshold, and a search operation has been performed based on the data index within a specified time period; and if the electronic device meets the third condition, the data index update operation is performed.

[0169] Furthermore, the search unit 902 is also used to determine, if it is not the first update, at least one of the following conditions: the time difference between the current time and the time of the last data index update operation is greater than a duration threshold; the number of times the update operation is executed within the current time period corresponding to the current time is less than a number threshold; the data to be updated this time is less than a specified quantity threshold; and the current time is in a nighttime period; and if the electronic device is determined to meet the fourth condition, the data index update operation is executed.

[0170] The fusion unit 903 is used to fuse the first search result and the second search result to obtain a fusion result.

[0171] Furthermore, the first search result includes multiple first matching data and the keyword matching degree corresponding to each first matching data, and the second search result includes multiple second matching data and the semantic matching degree corresponding to each second matching data. The fusion unit 903 is also used to determine the first weight corresponding to the first search result and the second weight corresponding to the second search result; based on the first weight and the second weight, the multiple first matching data and the second matching data are reordered to obtain a mixed ranking result, wherein the mixed ranking result includes multiple matching data in sequential order; the top N matching data in the mixed ranking result are taken as the fusion result, wherein N is an integer greater than 1.

[0172] Furthermore, the first weight and the second weight are the same.

[0173] Furthermore, the second weight is greater than the first weight.

[0174] Furthermore, the fusion unit 903 is also used to determine a reference matching degree based on the matching degree of at least one of the keywords in the first search results; if the reference matching degree is less than a preset threshold, then a second weight is set to be greater than the first weight; if the reference matching degree is greater than or equal to the preset threshold, then the first weight and the second weight are set to be the same.

[0175] Furthermore, the fusion unit 903 is also used to determine the maximum keyword matching degree in the first search result as a reference matching degree.

[0176] Processing unit 904 is used to obtain the target query result corresponding to the content to be queried based on the fusion result.

[0177] Furthermore, the processing unit 904 is also used to obtain the output result of the large language model based on the fusion result through the large language model, and to obtain the target query result corresponding to the content to be queried based on the output result.

[0178] Furthermore, the processing unit 904 is also used to send the fusion result to the cloud, trigger the cloud to generate prompt information based on the fusion result, input the prompt information into the large language model, and obtain the target query result corresponding to the query content based on the output result of the large language model.

[0179] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0180] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.

[0181] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0182] Please refer to Figure 10, which shows a structural block diagram of an electronic device provided in an embodiment of this application. The electronic device 100 can be a smartphone, tablet computer, e-reader, or other electronic device capable of running applications. The electronic device 100 in this application may include one or more of the following components: a processor 110, a memory 120, and one or more applications, wherein the one or more applications can be stored in the memory 120 and configured to be executed by one or more processors 110, and the one or more applications are configured to perform the methods described in the foregoing method embodiments.

[0183] Processor 110 may include one or more processing cores. Processor 110 connects to various parts within the electronic device 100 using various interfaces and lines, and performs various functions and processes data of the electronic device 100 by running or executing instructions, programs, code sets, or instruction sets stored in memory 120, and by calling data stored in memory 120. Optionally, processor 110 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 110 may integrate one or more of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 110 and may be implemented separately using a communication chip.

[0184] The memory 120 may include random access memory (RAM) or read-only memory (ROM). The memory 120 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data created by the electronic device 100 during use (such as phonebook data, audio and video data, chat log data, etc.).

[0185] Please refer to Figure 11, which shows a structural block diagram of a computer-readable medium provided in an embodiment of this application. The computer-readable medium 1100 stores program code that can be called by a processor to execute the methods described in the above method embodiments.

[0186] Computer-readable medium 1100 may be an electronic storage device such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, computer-readable medium 1100 includes non-transitory computer-readable storage medium. Computer-readable medium 1100 has storage space for program code 1110 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 1110 may be compressed, for example, in a suitable form.

[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A content query method, characterized in that, A search platform applied to an electronic device, the method comprising: Retrieve the content to be queried; In the business data of at least one target business accessed by the search platform, business data matching the content to be queried are searched based on keyword search and semantic search respectively, to obtain a first search result corresponding to the keyword search and a second search result corresponding to the semantic search. The first search result and the second search result are merged to obtain the fused result; Based on the fusion result, the target query result corresponding to the content to be queried is obtained.

2. The method according to claim 1, characterized in that, The step of searching for business data matching the query content in at least one target business data accessed by the search platform, based on keyword search and semantic search respectively, to obtain a first search result corresponding to the keyword search method and a second search result corresponding to the semantic search method, includes: Obtain a data index corresponding to the business data of at least one target business accessing the search platform, wherein the data index includes first index information and second index information, wherein the first index information includes preset keywords corresponding to the business data, and the second index information includes preset semantic features corresponding to the business data; Based on the keyword search method, business data matching the content to be queried is searched in the first index information to obtain the first search result; Based on the semantic search method, business data matching the content to be queried is searched in the second index information to obtain the second search result.

3. The method according to claim 2, characterized in that, The keyword search method includes a first keyword retrieval algorithm and a second keyword detection algorithm. Based on the keyword search method, the process of searching for business data matching the query content in the first index information to obtain the first search result includes: Based on the first keyword retrieval algorithm, business data matching the content to be queried is searched in the first index information to obtain the first matching result; Based on the second keyword retrieval algorithm, business data matching the content to be queried is searched in the first index information to obtain the second matching result; The first matching result and the second matching result are merged to obtain the first search result.

4. The method according to claim 3, characterized in that, The first keyword retrieval algorithm is an exact matching algorithm, and the second keyword retrieval algorithm is a fuzzy matching algorithm. In the first search result, the ranking priority of the second matching result is higher than that of the first matching result.

5. The method according to claim 2, characterized in that, The step of obtaining the data index corresponding to the business data of at least one target business accessing the search platform includes: Determine whether a data index has been established corresponding to the business data of at least one target business that is connected to the search platform; If established, obtain the data index corresponding to the business data of at least one target business accessing the search platform; If not established, establish a data index corresponding to the business data of at least one target business accessed by the search platform.

6. The method according to claim 2, characterized in that, Also includes: If a data index creation request is detected, determine whether the creation request is the first one. If this is the first time creating an index, then if the electronic device meets the first condition, the data index creation operation will be performed. If this is not the first time creating the data index, the data index creation operation is performed. If, during the data index creation operation, it is detected that the electronic device meets the second condition, the data index creation operation is interrupted.

7. The method according to claim 6, characterized in that, If this is the first creation, then if the electronic device meets the first condition, the data index creation operation is performed, including: If this is the first time creating the application, determine the target time period corresponding to the current moment; If the target time period is within the daytime period, the first condition is determined to include at least one of the following: the remaining power of the electronic device is greater than a first power threshold, the casing temperature of the electronic device is less than a first temperature threshold, the CPU load of the electronic device is less than a first load threshold, and the remaining storage space of the electronic device is greater than a first space threshold. If the target time period is within the nighttime period, the first condition is determined to include at least one of the following: the electronic device is in an idle state, the electronic device is in a screen-off state, the electronic device is in a charging state and the remaining power is greater than a second power threshold, the casing temperature of the electronic device is less than a first temperature threshold, the CPU load of the electronic device is less than a first load threshold, and the remaining storage space of the electronic device is greater than a first space threshold. Once the electronic device meets the first condition, the data index creation operation is performed.

8. The method according to claim 7, characterized in that, The daytime and nighttime time periods are set by analyzing user usage data of the electronic devices.

9. The method according to claim 6, characterized in that, If this is not the first time creating an index, then a data index creation operation is performed. During the data index creation operation, if it is detected that the electronic device meets the second condition, the current data index creation operation is interrupted, including: If this is not the first time it has been created, then determine the target time period corresponding to the current moment; If the target time period is within the daytime period, the second condition is determined to include the following: the casing temperature of the electronic device is greater than or equal to the first temperature threshold; the temperature difference of the casing temperature rise of the electronic device during the data index creation operation is greater than the second temperature threshold; or the power consumption of the electronic device during the data index creation operation exceeds the second power threshold. If the target time period is within the nighttime period, the second condition is determined to include the electronic device's casing temperature being greater than or equal to the first temperature threshold, the electronic device's power consumption exceeding the second power threshold during the data index creation operation, or the electronic device being in a screen-on state. The data index creation operation is executed, and if it is detected that the electronic device meets the second condition during the data index creation operation, the data index creation operation is interrupted.

10. The method according to claim 2, characterized in that, Also includes: If an update request for an established data index is detected, determine whether the update request is the first update. If this is the first update, then if the electronic device meets the third condition, the data index update operation will be performed. If this is not the first update, then if the electronic device meets the fourth condition, the data index update operation will be performed.

11. The method according to claim 10, characterized in that, If this is the first update, then if the electronic device meets the third condition, an update operation of the data index is performed, including: If this is the first update, the third condition is determined to include at least one of the following: the remaining battery power of the electronic device is greater than the second battery power threshold, the CPU load of the electronic device is less than the first load threshold, and a search operation has been performed based on the data index within the specified time period. If the electronic device meets the third condition, perform a data index update operation.

12. The method according to claim 10, characterized in that, If this is not the first update, then if the electronic device meets the fourth condition, an update operation of the data index is performed, including: If it is not the first update, the fourth condition is determined to include at least one of the following: the time difference between the current time and the time of the last data index update operation is greater than the duration threshold; the number of times the update operation is executed within the current time period corresponding to the current time is less than the number threshold; the data to be updated this time is less than the specified quantity threshold; and the current time is in the night time period. If the electronic device meets the fourth condition, perform a data index update operation.

13. The method according to any one of claims 1-12, characterized in that, The first search result includes multiple sequentially ordered first matching data, and the second search result includes multiple sequentially ordered second matching data. The step of merging the first and second search results to obtain a fusion result includes: Determine the first weight corresponding to the first search result and the second weight corresponding to the second search result; Based on the first weight and the second weight, multiple first matching data and second matching data are reordered to obtain a mixed sorting result, wherein the mixed sorting result includes multiple matching data ordered sequentially. The top N matching data in the mixed sorting results are taken as the fusion result, where N is an integer greater than 1.

14. The method according to claim 13, characterized in that, The first weight and the second weight are the same.

15. The method according to claim 13, characterized in that, The second weight is greater than the first weight.

16. The method according to claim 13, characterized in that, Determining the first weight corresponding to the first search result and the second weight corresponding to the second search result includes: A reference matching degree is determined based on the matching degree of at least one of the keywords in the first search results; If the reference matching degree is less than a preset threshold, then the second weight is set to be greater than the first weight; If the reference matching degree is greater than or equal to a preset threshold, then the first weight and the second weight are set to be the same.

17. The method according to claim 16, characterized in that, Determining a reference matching degree based on the matching degree of at least one of the keywords in the first search results includes: The highest keyword match score in the first search result is determined as the reference match score.

18. The method according to any one of claims 1-17, characterized in that, The process of obtaining the target query result corresponding to the content to be queried based on the fusion result includes: Using the large language model, based on the fusion result, the output result of the large language model is obtained, and based on the output result, the target query result corresponding to the content to be queried is obtained.

19. The method according to claim 18, characterized in that, The process involves obtaining the output of a large language model based on the fusion result, and then obtaining the target query result corresponding to the content to be queried based on the output result, including: The fusion result is sent to the cloud, triggering the cloud to generate a prompt message based on the fusion result. The prompt message is then input into a large language model, and the target query result corresponding to the query content is obtained based on the output of the large language model.

20. A content query device, characterized in that, A search platform for use within an electronic device, the device comprising: The retrieval unit is used to retrieve the content to be queried. The search unit is used to search for business data that matches the content to be queried in the business data of at least one target business accessed by the search platform, based on keyword search and semantic search respectively, and to obtain a first search result corresponding to the keyword search and a second search result corresponding to the semantic search. A fusion unit is used to fuse the first search result and the second search result to obtain a fusion result; The processing unit is used to obtain the target query result corresponding to the content to be queried based on the fusion result.

21. An electronic device, characterized in that, include: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to perform the method as described in any one of claims 1-19.

22. A computer-readable medium, characterized in that, The computer-readable medium stores processor-executable program code, which, when executed by the processor, causes the processor to perform the method according to any one of claims 1-19.