Indexing method and device for hot query

By building a three-level structure of geographical location, user portraits and query information, and using hierarchical clustering and large language models, the problems of large computing power consumption and delay in popular queries are solved, and fast response and high-quality query results are achieved.

CN120336383APending Publication Date: 2025-07-18WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510355547.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When handling popular queries, the model consumes a lot of computing power and the query results are delayed, making it difficult to respond quickly and improve the quality of answers.

Method used

By building a three-level structure of geographical location, user portraits and query information, using hierarchical clustering and large language models, pre-answer templates are constructed in advance, accurately locate user query content, and reduce model inference query needs.

Benefits of technology

It realizes rapid response and improves the quality of answers, reduces model resource consumption, and improves the accuracy and response speed of query results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336383A_ABST
    Figure CN120336383A_ABST
Patent Text Reader

Abstract

The invention provides a hot query indexing method and device. The hot query indexing method comprises the steps of obtaining query information input by a target user, a current position and user preference data; finding the ID of the location of the target user based on the distance between the current location and the stored location cluster; obtaining a user portrait corresponding to the user preference cluster with the highest similarity with the user preference data based on a similarity calculation result of the user preference data and the user preference clusters; finding a query information cluster to be queried based on the user location ID and the user portrait; calculating the similarity between the query information input by the target user and the query information cluster to be queried; and taking the pre-answer content corresponding to the query information cluster to be queried with the highest similarity as an output result. According to the invention, rapid answer response can be realized, and the answer quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tourism information retrieval, and particularly to an indexing method and device for popular queries. Background Art

[0002] For travel assistants, a large number of similar searches and similar answers are generated every day. In existing methods, the query statements input by users are usually converted into semantic vectors and key features to identify elements such as entities, relationships, and themes in the queries, as well as the logical relationships between them; then, based on the understood semantic information, retrieval is performed in the knowledge graph, database, or existing training data within the model; by calculating the similarity between the query vector and the stored knowledge vector, the information most relevant to the query is found; according to the retrieved relevant information, the large model uses its generation ability to organize and generate answers in the form of natural language. However, for popular query questions, if this query method is always adopted, it will greatly consume the computing power of the model, and when there are too many query questions at the same time, it will cause delays in the output query results. Summary of the Invention

[0003] The present invention provides an indexing method and device for popular queries, which are used to achieve fast answer response and improve the answer quality at the same time.

[0004] According to one aspect of the present invention, an indexing method for popular queries is provided, including:

[0005] Obtain the query information, current location, and user preference data input by the target user;

[0006] Find the ID of the location where the target user is located based on the distance between the current location and the stored location clusters;

[0007] Based on the similarity calculation result between the user preference data and the user preference clusters, obtain the user portrait corresponding to the user preference cluster with the highest similarity to the user preference data;

[0008] Find the query information cluster to be queried based on the user location ID and the user portrait;

[0009] Calculate the similarity between the query information input by the target user and the query information cluster to be queried;

[0010] Use the pre-answer content corresponding to the query information cluster to be queried with the highest similarity as the output result.

[0011] Optionally, before finding the ID of the location where the target user is located based on the distance between the current location and the stored location clusters, it includes:

[0012] The hierarchical clustering algorithm is used to divide the regions into multiple layers of location clusters, and the clustered location clusters are stored in the database.

[0013] Optionally, the step of using the hierarchical clustering algorithm to divide the regions into multiple layers of location clusters includes:

[0014] Performing hierarchical clustering on different regions according to three levels: city, administrative region, and popular points;

[0015] When the number of locations in the first region within different regions divided by the area of the first region is less than a preset threshold, the clustering algorithm based on points of interest is used to divide the points of interest data in the first region into different location clusters, obtaining multiple layers of location clusters.

[0016] Optionally, the step of finding the ID of the location where the target user is located based on the distance between the current location and the stored location clusters includes:

[0017] Converting the current location into a vector form;

[0018] Calculating the distances between the vector of the current location and the centers of multiple stored location clusters respectively, and the location corresponding to the location cluster with the closest distance is the ID of the location where the target user is located.

[0019] Optionally, before obtaining the user portrait corresponding to the user preference cluster with the highest similarity to the user preference data based on the similarity calculation result between the user preference data and the user preference clusters, it further includes:

[0020] Converting multiple user preference data of the same location collected into a vector form;

[0021] Clustering the vector forms of the user preference data to obtain multiple user preference clusters, and each user preference cluster corresponds to a user portrait.

[0022] Optionally, obtaining the user portrait corresponding to the user preference cluster with the highest similarity to the user preference data based on the similarity calculation result between the user preference data and the user preference clusters includes:

[0023] Converting the user preference data into a vector form;

[0024] Calculating the distances from the vector form of the user preference data to the centers of each user preference cluster respectively;

[0025] Obtaining the user portrait corresponding to the user preference cluster with the closest distance.

[0026] Optionally, before finding the query information cluster to be queried based on the ID of the user's location and the user portrait, it further includes:

[0027] Collect the query information of multiple users at the same location and with the same user profile;

[0028] Extract the semantic vectors of the query information of the users;

[0029] Cluster the semantic vectors of multiple query information to obtain multiple query information clusters.

[0030] Optionally, calculating the similarity between the query information input by the target user and the query information clusters to be queried includes:

[0031] Convert the query information input by the target user into a vector form;

[0032] Calculate the distances from the vector form of the query information to the centers of each query information cluster to be queried respectively, and the distance values reflect the similarity between the query information and the query information clusters to be queried.

[0033] Optionally, it further includes:

[0034] Use a large language model to summarize the pre-answers of each query information cluster and store the pre-answer content.

[0035] According to another aspect of the present invention, there is provided an indexing device for popular queries, including:

[0036] A data acquisition unit for acquiring the query information, current location, and user preference data input by the target user;

[0037] A location search unit for finding the ID of the location where the target user is located based on the distance between the current location and the stored location clusters;

[0038] A profile acquisition unit for acquiring the user profile corresponding to the user preference cluster with the highest similarity to the user preference data based on the similarity calculation result between the user preference data and the user preference clusters;

[0039] An information search unit for finding the query information clusters to be queried based on the user location ID and the user profile;

[0040] A similarity calculation unit for calculating the similarity between the query information input by the target user and the query information clusters to be queried;

[0041] A pre-answer output unit for using the pre-answer content corresponding to the query information cluster with the highest similarity as the output result.

[0042] According to another aspect of the present invention, there is provided an electronic device, and the electronic device includes:

[0043] At least one processor; and

[0044] A memory communicatively connected to the at least one processor; wherein,

[0045] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor can execute the indexing method for popular queries according to any embodiment of the present invention.

[0046] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the indexing method for popular queries according to any embodiment of the present invention when executed.

[0047] The technical solution of the embodiments of the present invention can accurately locate the content that the user wants to query by constructing a query method with a three-level structure of geographical location, user profile, and query information; in addition, by pre-constructing a pre-answer template for popular questions as the answer content to the user's query questions, it is not necessary for the large model to perform inference queries on the query information input by the user, enabling the model to quickly output query results.

[0048] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0050] Figure 1 is a flowchart of an indexing method for popular queries according to Embodiment 1 of the present invention;

[0051] Figure 2 is a flowchart of an indexing method for popular queries according to Embodiment 2 of the present invention;

[0052] Figure 3 is a structural diagram of an indexing device for popular queries according to Embodiment 3 of the present invention;

[0053] Figure 4 is a schematic structural diagram of an electronic device for implementing the indexing method for popular queries of the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0055] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0056] For travel assistants, a large number of similar searches and similar answers are generated every day. By clustering, the hot issues and hot topics of each type of user in each region are obtained, and an efficient hybrid index is constructed. When users ask some hot questions, they can rely on the efficient hybrid similarity index to quickly obtain standard pre-answers. A significant reduction in response time and an improvement in answer quality during travel will greatly enhance the user's travel experience.

[0057] Embodiment 1

[0058] Figure 1 The flowchart of an index method for popular queries is provided for Embodiment 1 of the present invention. As Figure 1 shown, the method includes:

[0059] S101. Obtain the query information, current location, and user preference data input by the target user.

[0060] Among them, the query information input by the user includes the query statement input by the user, such as "What are the delicious foods nearby" or "What are the interesting places nearby". The current location can be the location real-time located by the user terminal, or the location information input by the user. For example, the location A located by the mobile terminal, or the location B in the query statement "What are the delicious foods near location B" input by the user. The user's preference data includes various types of user preference data, such as travel time preference, user role preference, food preference, transportation preference, budget preference, travel companion preference, travel tag preference, etc.

[0061] S102. Find the ID of the location where the target user is located based on the distance between the current location and the stored location clusters.

[0062] Among them, the stored location clusters are location clusters generated by the system in advance through clustering geographical locations. These location clusters can be stored in a database. When a user needs to query information, the system can determine the similarity between the user's current location and the stored location clusters, and then determine the ID corresponding to the location cluster where the user is located.

[0063] It should be noted that for the same query information input at different locations, the query results may be different. For example, when a user inputs "What are the good places to eat nearby", the output query results will display different restaurants and foods according to the different current locations.

[0064] S103. Based on the similarity calculation result between the user preference data and the user preference clusters, obtain the user portrait corresponding to the user preference cluster with the highest similarity to the user preference data.

[0065] Among them, the stored user preference clusters are user preference clusters generated by the system in advance through clustering user preference data of different users. Each user preference cluster corresponds to a user portrait. The user portrait can be labels such as "student party", "office worker", "photography enthusiast", etc. The clustered user preference clusters can be stored in a database. When a user needs to query information, after determining the current location ID, the system can determine the similarity between the user's preference data and the user preference clusters, and then determine the user portrait.

[0066] It should be noted that for the same query information input by users, the query results may be different. For example, when a user inputs "What are the good places to eat nearby", the output query results will display different foods according to the different user preferences (preferring sweet or spicy foods).

[0067] S104. Find the query information clusters to be queried based on the user location ID and the user portrait.

[0068] After determining the user location ID and the user portrait, find multiple query information clusters to be queried. Each query information cluster corresponds to a query question, such as "food query", "location query", "route query", etc.

[0069] Query information of different users can be collected, the query information can be clustered, and the clustered query information clusters correspond to different popular questions. The clustered query information clusters can be stored in a database. When a user needs to query information, after determining multiple query information clusters to be queried, the query information of the user can be compared with the query information clusters to be queried for similarity determination, and then the category of the user's query information can be determined.

[0070] S105. Calculate the similarity between the query information input by the target user and the query information clusters to be queried.

[0071] After finding multiple query information clusters to be queried, subsequently, according to the calculation result of the similarity between the query information input by the target user and the query information clusters to be queried, find the query question corresponding to the query information cluster that the query information input by the user is closest to as the query question input by the target user.

[0072] It should be noted that for the same query information input by users, for different locations and users with different preferences, the query results may vary. For example, when a user inputs "What are the delicious foods nearby", the output query results will be different foods according to the user's current location and preferences.

[0073] S106. Use the pre-answer content corresponding to the query information cluster with the highest similarity as the output result.

[0074] For different query questions, when constructing a query index, the large model can be used to summarize each query question first, and the pre-answer of each query question is output, that is, each query information cluster corresponds to a pre-answer, and the pre-answer is stored. When the target user inputs query information, the query information is respectively calculated for similarity with multiple query information clusters to be queried, and then the query information cluster with the highest similarity is found, and its corresponding pre-answer can be used as the final output result.

[0075] The technical solution of the embodiment of the present invention can accurately locate the content that the user wants to query by constructing a query method with a three-level structure of geographical location, user portrait, and query information; in addition, by constructing a pre-answer template for popular questions in advance as the reply content to the user's query questions, it is not necessary for the large model to perform inference queries on the query information input by the user, so that the model can quickly output query results.

[0076] Embodiment 2

[0077] Figure 2 It is a flowchart of a method for indexing popular queries provided by Embodiment 2 of the present invention. As Figure 2 shown, the method further includes:

[0078] S201. Stratify and cluster different regions into three levels: city, administrative region, and popular location points.

[0079] S202. When the number of locations in the first region among different regions divided by the area of the first region is less than a preset threshold, then use a clustering algorithm based on points of interest to divide the point-of-interest data of the first region into different location clusters, obtaining multi-level location clusters.

[0080] In this embodiment, clustering can be performed according to the geographical location at three levels, including clustering at the city level, administrative region level, and popular location points. For example, the clustering levels are C City - D District - E Location (popular location). If the range of regional location points is large, for example, when the number of location points inside the region / the area of the region (the approximate area can be estimated with a square) is less than the threshold, then it can be shown that this region has the characteristic of "sparse location points over a large area". Therefore, more detailed density clustering can be performed on this basis to achieve better division. Then, the DBSCAN algorithm can be used to perform clustering based on points of interest on the basis of this region to improve the classification range accuracy. For example, the E location can be further clustered into: E Location: East E Area, West E Area, North E Area, etc.

[0081] Among them, the threshold can be set according to experience, or the machine learning idea of constructing: test set - randomly set the threshold - evaluate with the test set after clustering - adjust the threshold in reverse transmission can be adopted to obtain a relatively accurate threshold learned based on existing data.

[0082] S203. Convert multiple user preference data of the same location into vector form.

[0083] S204. Cluster the vector form of user preference data to obtain multiple user preference clusters, and each user preference cluster corresponds to a user portrait.

[0084] First, multiple user preference data of the same location can be collected, and the preference data of users can be encoded. One-hot encoding can be performed on different preference types of users. For example, for different preferences of users: travel time preference, user role preference, food preference, transportation preference, budget preference, travel companion preference, travel label preference, etc., the preference data of each user can be converted into vector representation, and each dimension of preference in the vector uses a specific label.

[0085] The kmeans clustering can be used to cluster the preference data of users, determine the number of clusters k, randomly select k points as the initial centers, calculate the distance from users to the centers, assign users to categories, and iteratively adjust the center positions. Eventually, different user preference clusters and each user portrait are formed to provide personalized travel recommendations.

[0086] After clustering is completed, users with similar personality preferences can be grouped together to form different user portraits. These user portraits can better understand the needs of users, thereby providing more personalized travel recommendations.

[0087] By combining one-hot encoding and clustering, users can be divided into different clusters according to different user preference data. For example, the finally divided user preference clusters can be labels such as student groups, office workers, photography enthusiasts, etc., or more complex and diverse clustering results.

[0088] S205. Collect the query information of multiple users with the same location and the same user portrait; extract the semantic vectors of the users' query information.

[0089] S206. Cluster the semantic vectors of multiple query information to obtain multiple query information clusters.

[0090] The bge model or other Text-Embedding models can be used to extract the semantic feature vectors of query information. It can capture the relationship between the front and back information in the text, thereby better understanding the semantics of the text. By inputting the question text of the query information into the bge model, a vector that can represent the semantics of the question can be obtained.

[0091] Then perform clustering analysis on the extracted semantic feature vectors. The density-based DBSCAN clustering algorithm can be selected. The DBSCAN algorithm is a density-based clustering method that does not require pre-specifying the number of clusters and can automatically identify clusters with different densities. This method is very suitable for processing the clustering of problem semantic vectors because it can identify semantically similar problems even if they are different in text expression.

[0092] For example, when there are the following query information: A: What are the delicious foods in Wuhan University? B: What are the interesting spots in Wuhan University? C: Which restaurants in Wuhan University should be avoided? By extracting the semantic vectors of texts A, B, and C and calculating the similarity between the semantic vectors, it can be obtained that query information A and C are closer, so query information A and C can be clustered into one cluster.

[0093] S207. Obtain the query information, current location, and user preference data input by the target user.

[0094] S208. Convert the current location, user preference data, and input query information of the target user into vector form.

[0095] It should be noted that the current location, user preference data, and input query information of the target user can be respectively converted into vector forms for subsequent similarity comparison, so as to accurately find the query problem type of the target user.

[0096] S209. Calculate the distances between the vector of the current location and the centers of multiple stored location clusters respectively, and the location corresponding to the location cluster with the closest distance is the ID of the location where the target user is located.

[0097] The similarity between the location cluster of the three-level hierarchical geographical location clustering and the user's current location can be calculated, that is, calculate the distance between the user's current location and the center of the location cluster, and the location corresponding to the location cluster with the closest distance is the ID of the location where the target user is located. When there are sub-locations in the current location where the target user is located, the Euclidean distance based on the points of interest can be used to calculate the distance from the current location to the center points of each sub-interval, and then the point with the closest distance is found as the sub-interval where the target user is located.

[0098] S210. Calculate the distances from the vector form of the user preference data to the centers of each user preference cluster respectively.

[0099] S211. Obtain the user portrait corresponding to the user preference cluster with the closest distance.

[0100] After determining the ID of the user's location, the similarity between the user's preference data and the user preference clusters can be judged, that is, calculate the distances from the vector form of the user preference data to the centers of each user preference cluster, and the user portrait corresponding to the user preference cluster with the closest distance.

[0101] S212. Calculate the distances from the vector form of the query information to the centers of each query information cluster to be queried respectively, and the distance value reflects the similarity between the query information and the query information cluster to be queried.

[0102] After determining the ID of the user's location and the user portrait, find multiple query information clusters to be queried, and judge the similarity between the user's query information and the query information clusters to be queried, that is, calculate the distances from the vector form of the query information to the centers of each query information cluster to be queried, and the query problem corresponding to the query information cluster with the closest distance value is the category of the user's query information.

[0103] S213. Use the pre-answer content corresponding to the query information cluster to be queried with the highest similarity as the output result.

[0104] In one embodiment, it further includes using a large language model to summarize the pre-answer of each query information cluster and storing the pre-answer content.

[0105] Specifically, based on the established three-layer clustering, query information clusters with similar characteristics can be obtained. When constructing the index, each query information cluster is summarized using an LLM (Large Language Model), and a preliminary answer for each cluster (query statement) can be obtained and stored as a pre-answer. For example, for the query statement: Food recommendations in place A, a preliminary answer can be obtained first as a pre-answer.

[0106] By summarizing similar questions once as standard answers, subsequent answers can be quickly output in the form of pre-answers, without the need to query and retrieve using the large model for each question query, thus significantly reducing resource requirements.

[0107] The technical solution of the embodiment of the present invention can more accurately identify and recommend popular points related to the user's location through geographical location clustering, thereby enhancing the user experience in the cultural and tourism scenario. Combining with the DBSCAN algorithm, more detailed clustering can be performed on a larger range of spots, improving the accuracy and practicality of recommendations. Through personality preference clustering (traveler portrait), a unique user portrait can be constructed based on the user's travel time, role, preferences, etc., to achieve personalized recommendations. Through Kmeans clustering, users with similar preferences can be grouped together to provide travel suggestions that better meet their needs: Through question semantic vector clustering, the beg model is used to extract the semantic features of long texts, which can more accurately understand the essence of the user's question. The density-based DBSCAN clustering algorithm can automatically identify semantically similar questions, even if they are different in text expression, thus improving the accuracy of recommendations. Finally, through the extraction and clustering of the feature vectors of the user's questions, the system can provide more accurate pre-answer content and improve user satisfaction.

[0108] Embodiment III

[0109] Figure 3 It is a schematic structural diagram of a popular query index device provided in Embodiment III of the present invention. As Figure 3 shown, the device includes:

[0110] A data acquisition unit 301, configured to acquire query information, the current location, and user preference data input by a target user;

[0111] A location search unit 302, configured to find the ID of the location where the target user is located based on the distance between the current location and the stored location clusters;

[0112] A portrait acquisition unit 303, configured to acquire a user portrait corresponding to the user preference cluster with the highest similarity to the user preference data based on the similarity calculation result between the user preference data and the user preference clusters;

[0113] An information search unit 304, configured to find a query information cluster to be queried based on the user location ID and the user profile;

[0114] A similarity calculation unit 305, configured to calculate the similarity between the query information input by the target user and the query information cluster to be queried;

[0115] A pre-answer output unit 306, configured to use the pre-answer content corresponding to the query information cluster with the highest similarity as the output result.

[0116] The index device for popular queries provided by the embodiments of the present invention can execute the index device for popular queries provided by any embodiment of the present invention, and has function modules and beneficial effects corresponding to the execution method.

[0117] Embodiment 4

[0118] Figure 4 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0119] As Figure 4 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0120] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0121] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as an indexing method for popular queries.

[0122] In some embodiments, an indexing method for popular queries can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the indexing method for popular queries described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute an indexing method for popular queries in any other suitable manner (e.g., by means of firmware).

[0123] The various embodiments of the systems and technologies described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0124] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0125] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0126] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0127] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0128] A computing system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0129] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0130] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An indexing method for popular queries, characterized in that, Including: Obtain the query information, current location, and user preference data input by the target user; Find the ID of the location where the target user is located based on the distance between the current location and the stored location clusters; Based on the similarity calculation result between the user preference data and the user preference clusters, obtain the user portrait corresponding to the user preference cluster with the highest similarity to the user preference data; Find the query information cluster to be queried based on the user location ID and the user portrait; Calculate the similarity between the query information input by the target user and the query information cluster to be queried; Use the pre-answer content corresponding to the query information cluster to be queried with the highest similarity as the output result.

2. The indexing method for popular queries according to claim 1, characterized in that, Before the step of finding the ID of the location where the target user is located based on the distance between the current location and the stored location clusters, it includes: Use the hierarchical clustering algorithm to divide the region into multiple layers of location clusters and store the clustered location clusters in the database.

3. The indexing method for popular queries according to claim 2, characterized in that, The step of using the hierarchical clustering algorithm to divide the region into multiple layers of location clusters includes: Perform hierarchical clustering on different regions according to three levels: city, administrative region, and popular location; When the number of locations in the first region in different regions divided by the area of the first region is less than the preset threshold, use the clustering algorithm based on points of interest to divide the points of interest data in the first region into different location clusters to obtain multiple layers of location clusters.

4. The indexing method for popular queries according to claim 3, wherein The step of finding the ID of the location where the target user is located based on the distance between the current location and the stored location clusters includes: Convert the current location into a vector form; Calculate the distances between the vector of the current location and the centers of multiple stored location clusters respectively, and the location corresponding to the location cluster with the closest distance is the ID of the location where the target user is located.

5. The indexing method for popular queries according to claim 1, characterized in that, Before the step of obtaining the user portrait corresponding to the user preference cluster with the highest similarity to the user preference data based on the similarity calculation result between the user preference data and the user preference clusters, it also includes: Convert multiple user preference data collected at the same location into vector form; Cluster the vector forms of the user preference data to obtain multiple user preference clusters, and each user preference cluster corresponds to a user portrait.

6. The indexing method for popular queries according to claim 1, wherein Based on the similarity calculation result between the user preference data and the user preference clusters, the step of obtaining the user portrait corresponding to the user preference cluster with the highest similarity to the user preference data includes: Convert the user preference data into vector form; Calculate the distances from the vector form of the user preference data to the centers of each user preference cluster respectively; Obtain the user portrait corresponding to the user preference cluster with the closest distance.

7. The indexing method for popular queries according to claim 1, characterized in that, Before the step of finding the query information cluster to be queried based on the user location ID and the user portrait, it also includes: Collect the query information of multiple users at the same location and with the same user portrait; Extract the semantic vectors of the query information of the users; Perform clustering processing on the semantic vectors of multiple query information to obtain multiple query information clusters.

8. The indexing method for popular queries according to claim 7, characterized in that, The step of calculating the similarity between the query information input by the target user and the query information cluster to be queried includes: Convert the query information input by the target user into vector form; Calculate the distances from the vector form of the query information to the centers of each query information cluster to be queried respectively. The distance values reflect the similarity between the query information and the query information clusters to be queried.

9. The indexing method for popular queries according to claim 7, characterized in that, Further comprising: Using a large language model to summarize the pre-answers of each query information cluster and storing the pre-answer content.

10. An indexing device for popular queries, characterized in that, Comprising: A data acquisition unit for acquiring query information, the current location, and user preference data input by the target user; A location search unit for finding the ID of the location where the target user is located based on the distance between the current location and the stored location clusters; A portrait acquisition unit for acquiring the user portrait corresponding to the user preference cluster with the highest similarity to the user preference data based on the similarity calculation result between the user preference data and the user preference clusters; An information search unit for finding the query information clusters to be queried based on the user location ID and the user portrait; A similarity calculation unit for calculating the similarity between the query information input by the target user and the query information clusters to be queried; A pre-answer output unit for using the pre-answer content corresponding to the query information cluster to be queried with the highest similarity as the output result.