Artificial intelligence-based data query method and device, computer device and medium

CN122817296APending Publication Date: 2026-09-25CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610804898.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本申请实施例的目的在于提出一种基于人工智能的数据查询方法、装置、计算机设备及存储介质,以解决现有的搜索推荐系统存在数据查询效率较低的技术问题

Benefits of technology

[0010]上述基于人工智能的数据查询方法、装置、计算机设备及存储介质所实现的方案中,首先接收用户在搜索框输入的前缀数据;然后基于预设的语义生成器对所述前缀数据进行语义提取处理,得到对应的语义数据;之后基于预设的聚类搜索策略对所述语义数据进行检索处理,得到对应的候选查询列表;后续获取与所述用户对应的历史查询序列,以及获取所述用户的画像数据;进一步基于预设的编码器对所述前缀数据、所述候选查询列表、所述历史查询序列以及所述画像数据进行融合处理,得到对应的上下文表征;并基于预设的解码器对所述上下文表征进行自回归逐步生成处理,得到对应的查询列表;最后基于预设的后处理策略对所述查询列表进行优化处理得到对应的目标查询列表,并对所述目标查询列表进行输出。不同于现有的基于多阶段级联的流水线架构的数据查询方法,基于以上的自动化处理流程,本申请通过采用生成式架构来替代多阶段级联框架,包括前缀语义提取、聚类搜索、编码器多源融合和解码器搜索生成的整个数据查询流程,均在单次模型推理中完成,从而消除了多阶段级联带来的管道延迟和模块间通信开销,端到端延迟控制在毫秒级别,能够在毫秒级别自动准确地直接生成高质量的目标查询列表,有效地提高了数据查询的处理效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817296A_ABST
    Figure CN122817296A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of artificial intelligence, and relates to a data query method and device based on artificial intelligence, a computer device and a storage medium, comprising the following steps: receiving prefix data input by a user in a search box; performing semantic extraction processing on the prefix data based on a semantic generator to obtain semantic data; performing retrieval processing on the semantic data based on a clustering search strategy to obtain a candidate query list; obtaining a historical query sequence corresponding to the user and portrait data of the user; performing fusion processing on the prefix data, the candidate query list, the historical query sequence and the portrait data based on an encoder to obtain context representation; performing autoregressive step-by-step generation processing on the context representation based on a decoder to obtain a query list; and performing optimization processing on the query list based on a post-processing strategy to obtain a target query list and output the target query list. The application can be applied to data query business scenarios in the fields of financial technology and medical health, and effectively improves the processing efficiency of data query.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology and can be applied to fields such as fintech and healthcare, particularly to data query methods, devices, computer equipment, and storage media based on artificial intelligence. Background Technology

[0002] Existing search recommendation systems, especially when handling query completion and search suggestion tasks, generally employ a multi-stage cascading pipeline architecture. This architecture typically includes several independent stages, such as recall, coarse ranking, and re-ranking. In the recall stage, the system quickly filters a small number of entries related to the user's input prefix from a massive pool of candidates, often using techniques such as inverted indexes and vector similarity retrieval. In the subsequent ranking stage, a more complex but computationally expensive model is used to refine the ranking of the recall results to determine the final recommendation list presented to the user. This multi-stage processing flow means that data needs to be serialized, transferred, and waited for between different modules. Each stage requires independent computation and I / O overhead, and the data serialization / deserialization and network transmission between stages introduce significant cumulative latency. In scenarios with extremely high real-time requirements for user interaction (such as input-based recommendations), this cascading latency severely impacts the smoothness of the user experience, resulting in low data query processing efficiency.

[0003] For example, in the financial insurance sector, when a user enters "critical illness insurance" in the search box, the traditional multi-stage architecture requires first recalling candidates through an inverted index, then filtering through a coarse-ranking model, and finally sorting through a fine-ranking model. The entire process typically has a latency of over 100 milliseconds, which users will clearly perceive as a lag in the recommended list when typing quickly. Similarly, in the healthcare sector, when a user enters "headache relief," the system needs to sequentially complete vector retrieval, coarse-ranking filtering, and fine-ranking scoring. This multi-stage sequential processing makes it impossible for the recommended results to keep pace with the user's input, and the latency significantly degrades the search experience, especially when users are eager to obtain medical guidance.

[0004] Therefore, there is an urgent need to provide an intelligent search and recommendation method to eliminate the cumulative delay caused by multi-stage cascading and improve the data query efficiency in real-time interactive scenarios. Summary of the Invention

[0005] The purpose of this application is to propose a data query method, apparatus, computer device, and storage medium based on artificial intelligence, so as to solve the technical problem of low data query efficiency in existing search and recommendation systems.

[0006] Firstly, an artificial intelligence-based data query method is provided, including: Receive prefix data entered by the user in the search box; The prefix data is semantically extracted based on a preset semantic generator to obtain the corresponding semantic data. The semantic data is retrieved and processed based on a preset clustering search strategy to obtain a corresponding candidate query list; Obtain the historical query sequence corresponding to the user, and obtain the user's profile data; The prefix data, the candidate query list, the historical query sequence, and the profile data are fused based on a preset encoder to obtain the corresponding context representation; The context representation is generated stepwise by autoregression based on a preset decoder to obtain the corresponding query list; The query list is optimized based on a preset post-processing strategy to obtain the corresponding target query list, and the target query list is output.

[0007] Secondly, an artificial intelligence-based data query device is provided, including: The receiving module is used to receive prefix data entered by the user in the search box; The extraction module is used to perform semantic extraction processing on the prefix data based on a preset semantic generator to obtain the corresponding semantic data; The retrieval module is used to perform retrieval processing on the semantic data based on a preset clustering search strategy to obtain the corresponding candidate query list; The acquisition module is used to acquire the historical query sequence corresponding to the user, and to acquire the user's profile data; The fusion module is used to perform fusion processing on the prefix data, the candidate query list, the historical query sequence and the profile data based on a preset encoder to obtain the corresponding context representation; The generation module is used to perform autoregressive stepwise generation processing on the context representation based on a preset decoder to obtain the corresponding query list; The processing module is used to optimize the query list based on a preset post-processing strategy to obtain the corresponding target query list, and output the target query list.

[0008] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described artificial intelligence-based data query method.

[0009] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the aforementioned data query method based on artificial intelligence.

[0010] In the aforementioned scheme implemented by the AI-based data query method, apparatus, computer device, and storage medium, the following steps are taken: First, prefix data input by the user in the search box is received; then, semantic extraction processing is performed on the prefix data based on a preset semantic generator to obtain corresponding semantic data; subsequently, the semantic data is retrieved based on a preset clustering search strategy to obtain a corresponding candidate query list; next, historical query sequences corresponding to the user and user profile data are obtained; further, the prefix data, candidate query list, historical query sequence, and profile data are fused based on a preset encoder to obtain a corresponding context representation; and the context representation is autoregressively generated based on a preset decoder to obtain a corresponding query list; finally, the query list is optimized based on a preset post-processing strategy to obtain a corresponding target query list, and the target query list is output. Unlike existing data query methods based on multi-stage cascaded pipeline architectures, this application replaces the multi-stage cascaded framework with a generative architecture based on the above automated processing flow. The entire data query process, including prefix semantic extraction, cluster search, encoder multi-source fusion, and decoder search generation, is completed in a single model inference, thereby eliminating pipeline latency and inter-module communication overhead caused by multi-stage cascading. End-to-end latency is controlled at the millisecond level, enabling the automatic and accurate direct generation of high-quality target query lists at the millisecond level, effectively improving the processing efficiency of data query. Attached Figure Description

[0011] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of the artificial intelligence-based data query method according to this application; Figure 3 This is a schematic diagram of a structure of an embodiment of the artificial intelligence-based data query device according to this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0014] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0016] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0017] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0018] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0019] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0020] It should be noted that the AI-based data query method provided in this application is generally executed by a server / terminal device, and correspondingly, the AI-based data query device is generally located in the server / terminal device.

[0021] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0022] Continue to refer to Figure 2 The flowchart illustrates an embodiment of the AI-based data query method according to this application. The order of the steps in the flowchart can be changed, and some steps can be omitted, depending on different needs. The AI-based data query method includes the following steps: Step S201: Receive the prefix data entered by the user in the search box.

[0023] In this embodiment, the artificial intelligence-based data query method runs on an electronic device (e.g., Figure 1 The server / terminal device shown can obtain the prefix data entered by the user in the search box via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future wireless connection methods. The implementing entity of this application is specifically a data query system, which can be simply referred to as the system. The aforementioned prefix data can be the prefix entered by the user in the system's search box according to their actual query needs.

[0024] This application can be applied to data query scenarios in the financial insurance and healthcare sectors. For example, in a financial insurance product comparison scenario, a user enters the prefix "which critical illness insurance is best" in the search box of a financial app, and the system receives this prefix p = "which critical illness insurance is best" as input. In this case, the prefix is ​​semantically ambiguous; "critical illness insurance" could refer to different product categories such as single-payment, multiple-payment, term, and return-of-premium insurance, while "which is best" indicates that the user has a comparison intention but has not yet specified the comparison dimensions. Alternatively, in a financial insurance product selection scenario, a user enters the prefix "large-amount deposit" in the search box of a bank app, and the system receives this prefix p = "large-amount deposit" as input. This prefix is ​​extremely short and semantically sparse, potentially pointing to multiple completely different query intentions such as "large-amount deposit interest rate," "large-amount deposit bank," and "minimum deposit amount for large-amount deposit."

[0025] Furthermore, in symptom search scenarios within the healthcare field, when a user enters the prefix "headache hang what" in the search box of a health app, the system receives this prefix p = "headache hang what" as input. This prefix is ​​semantically incomplete; "hang what" could refer to different intentions such as "which department to consult," "which appointment to register for," or "which specialist appointment to register for." Similarly, in drug search scenarios within the healthcare field, when a user enters the prefix "amoxicillin" in the search box of a pharmaceutical e-commerce platform, the system receives this prefix p = "amoxicillin" as input. This prefix is ​​an incomplete input of the drug name and could lead to multiple search directions such as "amoxicillin instructions," "amoxicillin side effects," "amoxicillin dosage for children," or "differences between amoxicillin and cephalosporins."

[0026] Step S202: Based on a preset semantic generator, perform semantic extraction processing on the prefix data to obtain the corresponding semantic data.

[0027] In this embodiment, the specific implementation process of performing semantic extraction processing on the prefix data based on the preset semantic generator to obtain the corresponding semantic data will be further described in detail in subsequent specific embodiments, and will not be elaborated on here.

[0028] Step S203: The semantic data is retrieved based on a preset clustering search strategy to obtain the corresponding candidate query list.

[0029] In this embodiment, the specific implementation process of retrieving and processing the semantic data based on the preset clustering search strategy to obtain the corresponding candidate query list will be described in more detail in subsequent specific embodiments, and will not be elaborated on here.

[0030] Step S204: Obtain the historical query sequence corresponding to the user, and obtain the user's profile data.

[0031] In this embodiment, the aforementioned historical query sequence originates from the user behavior log database in the system backend. This database continuously records all historical interaction behaviors of each user in financial and medical search scenarios. Each record includes: a unique user identifier (User ID), a search timestamp, the entered query text, and subsequent behaviors for that search (which result was clicked, whether it was converted into a transaction, etc.).

[0032] The extraction process includes: When a user enters the current prefix 'p' in the search box, the system uses the user's User ID as the key to query the user's past N search records (N is usually 5 to 15) from the behavior log database. The query results are sorted in reverse chronological order, with the most recent search at the top and earlier searches following in sequence. For example, if the user's three most recent searches are "iPhone 16 price", "AirPods Pro review", and "fund investment strategy", the extracted results would be ["iPhone 16 price", "AirPods Pro review", "fund investment strategy"]. Then, the extracted N historical queries are arranged in reverse chronological order, separated by spaces, forming a continuous text segment. This text segment is H_u (historical query sequence), which is directly used as part of the input sequence in subsequent encoding. It is important to note that if a historical query contains sensitive information or has been deleted by the user, that record is automatically skipped during extraction to ensure the compliance of the input sequence. Meanwhile, if there are fewer than N historical records (such as new users), padding is performed using a special padding marker to ensure that the sequence length remains consistent.

[0033] The user profile data mentioned above comes from the system's User Profile Database, from which the user's profile data can be extracted. This database stores structured feature tags for each user, which are aggregated from data from multiple channels: first, information actively filled in by the user during registration (such as age, occupation, and city); second, behavioral data generated by the user during product use (such as frequently searched categories, click preferences, and portfolio information); and third, anonymized user features provided by third-party data providers (such as risk preference level and asset size range). User profiles are stored in key-value pair format. For example, a user's profile might contain the following structured tags: {Interest Areas: Technology, Category Focused On: Consumer Electronics, Investment Preference: Conservative, Asset Size: Medium, City: Shenzhen}. These tags are discrete, non-continuous categorical features and cannot be directly used as a text input encoder.

[0034] The system pre-sets a set of natural language templates to convert structured tags into continuous natural language descriptions using fixed sentence patterns. The conversion rule is as follows: following the sentence pattern "This user is an enthusiast of [interest area], follows [category of interests], has a [financial preference], possesses [asset size], and resides in [city]," the values ​​of each tag are sequentially filled into the corresponding slots. For example, using the above profile as an example, the converted result is: "This user is a technology enthusiast, follows the consumer electronics field, has a conservative financial preference, possesses a medium-sized asset size, and resides in Shenzhen."

[0035] Step S205: Based on a preset encoder, the prefix data, the candidate query list, the historical query sequence, and the profile data are fused to obtain the corresponding context representation.

[0036] In this embodiment, the specific implementation process of fusing the prefix data, the candidate query list, the historical query sequence, and the profile data based on the preset encoder to obtain the corresponding context representation will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0037] Step S206: Based on a preset decoder, perform autoregressive stepwise generation processing on the context representation to obtain the corresponding query list.

[0038] In this embodiment, the specific implementation process of performing autoregressive stepwise generation processing on the context representation based on a preset decoder to obtain the corresponding query list will be described in more detail in subsequent specific embodiments, and will not be elaborated on here.

[0039] Step S207: Optimize the query list based on a preset post-processing strategy to obtain the corresponding target query list, and output the target query list.

[0040] In this embodiment, the specific implementation process of optimizing the query list based on the preset post-processing strategy to obtain the corresponding target query list and outputting the target query list will be described in more detail in subsequent specific embodiments of this application, and will not be elaborated on here. Alternatively, the output processing of the target query list can be completed by presenting it to the user in the form of completion suggestions.

[0041] Unlike existing data query methods based on multi-stage cascaded pipeline architectures, this application replaces the multi-stage cascaded framework with a generative architecture based on the above automated processing flow. The entire data query process, including prefix semantic extraction, cluster search, encoder multi-source fusion, and decoder search generation, is completed in a single model inference, thereby eliminating pipeline latency and inter-module communication overhead caused by multi-stage cascading. End-to-end latency is controlled at the millisecond level, enabling the automatic and accurate direct generation of high-quality target query lists at the millisecond level, effectively improving the processing efficiency of data query.

[0042] In some alternative implementations, step S202 includes the following steps: Step S2021: Perform word segmentation on the prefix data to obtain the corresponding word sequence.

[0043] In this embodiment, when a user enters the prefix data p (such as "smart") in the search box, the system first calls the lexer to perform word segmentation on the prefix, dividing it into a lexer sequence {tok_1, tok_2, ..., tok_n}.

[0044] Step S2022: Obtain a preset designated marker and insert the designated marker into a designated position in the lexical sequence to obtain the corresponding lexical embedding sequence.

[0045] In this embodiment, the specified markers include a classification start marker and a separator marker. By inserting a classification start marker `t_cls` at the beginning of the lexical sequence and a separator marker `t_sep` at the end of the sequence according to a unified input format protocol, a structured input sequence `x = {t_cls, tok_1, tok_2, ..., tok_n, t_sep}` is constructed, which is the lexical embedding sequence. This lexical embedding sequence is the sole input to the RQ-VAE encoder, where `t_cls` identifies the input start point and guides the encoder to extract global semantics, and `t_sep` identifies prefix boundaries to prevent the encoder from confusing subsequent search results with prefixes.

[0046] Step S2023: Invoke the pre-built semantic generator.

[0047] In this embodiment, the semantic generator can specifically be an RQ-VAE encoder that has been trained and whose parameters have been frozen. The RQ-VAE encoder is composed of m layers of residual quantization modules connected in series (m is usually 3 to 5 layers).

[0048] Step S2024: Perform forward computation processing on the word embedding sequence based on the semantic generator to obtain the corresponding discrete semantic representation.

[0049] In this embodiment, after receiving the above-mentioned lexical embedding sequence, the semantic generator encodes the embedding layer by layer through a multi-layer residual quantization module. The specific implementation process includes: Layer 0 (coarsest granularity) quantization: The semantic generator receives H_0 in Layer 0 and maps it to a continuous vector r_0 ∈ Rd through a feedforward network. Then, it searches for the centroid vector e_{c_0} with the closest Euclidean distance to r_0 in the Layer 0 codebook C_0 = {e_01, e_02, ..., e_0{K_0}}, i.e., c_0 = argmin_j ‖r_0 - e_0j‖². At the same time, it calculates the quantization residual z_0 = r_0 - e_{c_0} and cuts off the gradient backpropagation of the centroid vector by stopping the gradient operation sg[·] to ensure that the quantization operation does not affect the update of the codebook. The residual z_0 is then fed into Layer 1.

[0050] Quantization from layer 1 to layer m-2: The i-th layer (1 ≤ i ≤ m-2) receives the residual z_{i-1} from the previous layer, maps it to a continuous vector r_i through the feedforward network, finds the nearest centroid e_{c_i} in the codebook C_i of the i-th layer, and calculates the residual z_i = r_i - e_{c_i}, also using the stopping gradient operation. The residual z_i continues to be passed to the next layer.

[0051] The (m-1)th layer (fineest granularity) quantization: The last layer receives the residual z_{m-2}, maps it to r_{m-1}, and searches for the nearest centroid e_{c_{m-1}} in the codebook C_{m-1} to obtain the final quantization code c_{m-1}. This layer no longer calculates the residual.

[0052] After m layers of quantization, the prefix p is represented as a set of hierarchical discrete codes ID(p) = {c_0, c_1, ..., c_{m-1}}. Among them, c_0 encodes the coarsest granular semantics (such as "technology"), c_1 encodes the medium granular semantics (such as "electronic products"), and c_{m-1} encodes the finest granular semantics (such as "smart devices").

[0053] The RQ-VAE encoder outputs hierarchical semantic IDs (i.e., discrete semantic representations), ID(p) = {c_0, c_1,..., c_{m-1}}, which are output as a sequence of integer indices. Each c_i corresponds to the centroid index in its respective layer's codebook, with a value range of [0, K_i - 1]. This index sequence is directly used as the key for subsequent retrieval steps. Since all parameters of the RQ-VAE (including the embedding table, feedforward network weights, and codebook centroids) are frozen during the inference phase, for the same prefix p, the semantic ID sequence output by the encoder is completely consistent regardless of when it is input, with no randomness or drift. This determinism ensures long-term consistency between the offline-built semantic codebook index and online retrieval, enabling the reliable execution of the fine-to-coarse retrieval strategy.

[0054] Step S2025: The discrete semantic representation is used as the semantic data.

[0055] Based on the above processing flow, the prefix encoding and semantic ID extraction provided in this application are essentially a hierarchical discretization process. Short prefixes are semantically sparse and have extremely low discriminative power in the original text space, making them unsuitable for efficient retrieval. The RQ-VAE encoder, through layer-by-layer residual quantization, gradually approximates continuous semantic vectors into a set of discrete codes: each layer captures semantic information at a specific granularity, with the coarse-grained layer determining the semantic class and the fine-grained layer determining the specific subclass. This hierarchical design enables the semantic representation to possess both compressibility (compressing a high-dimensional vector into m integer indices) and hierarchical structure (supporting multi-level matching from coarse to fine). Furthermore, the parameter freezing mechanism ensures the stability and reproducibility of the encoding process during the inference phase, providing reliable index keys for subsequent retrieval steps.

[0056] In some optional implementations of this embodiment, step S203 includes the following steps: Step S2031: Invoke the pre-built offline index.

[0057] In this embodiment, the construction process of the offline index includes: in the preprocessing stage before the system goes online, all high-quality queries in the historical logs are batch-encoded using an RQ-VAE (semantic generator) with frozen parameters. For each historical query q, the same RQ-VAE forward calculation process as in the inference stage is executed to obtain its hierarchical semantic ID, ID(q) = {c_0, c_1, ..., c_{m-1}}. Subsequently, the original text q of each query, its complete semantic ID sequence ID(q), and the historical statistical information of the query (such as click-through rate and conversion rate) are stored in the inverted index as the corresponding offline index. The primary key of the index is the hierarchical semantic ID, and each ID level corresponds to an index field. For example, the level 0 code c_0 corresponds to the first-level index, the level 1 code c_1 corresponds to the second-level index, and so on. The index structure supports multi-level queries: it can perform precise searches using the complete ID sequence, or range searches using only one or several levels of code. This offline index remains static after the system goes live, and is only incrementally rebuilt during periodic data updates.

[0058] Step S2032: Based on the offline index, perform precise matching and retrieval of the semantic data to obtain the corresponding first candidate query result.

[0059] In this embodiment, the semantic ID obtained by prefix encoding (i.e., the semantic data mentioned above) is used as the retrieval key value to perform an exact match query in the offline index. The query condition requires that the semantic ID of the candidate query q is completely consistent with ID(p) at all levels, that is, ID(q) = ID(p), which is equivalent to requiring For each i ∈ [0, m-1], c_i(q) = c_i(p). The index engine directly locates the corresponding hash bucket based on this condition, with a time complexity of approximately constant. All successfully matched queries constitute the exact match set S_fine (i.e., the first candidate query result). Queries in this set are semantically closest to the prefix because their hierarchical semantic IDs are consistent at every granularity level. For example, if the semantic ID of the prefix "smart" is {technology, electronic products, smart devices}, then S_fine contains all historical queries with the exact same semantic ID, such as "smartphone latest price" and "smart watch battery life test".

[0060] Step S2033: Perform coarse-grained extended retrieval of the semantic data based on the offline index to obtain the corresponding second candidate query results.

[0061] In this embodiment, when the number of candidates for the first candidate query result, i.e., S_fine, is insufficient to meet a preset threshold, or when the system needs to increase candidate diversity, a hierarchical expansion retrieval from fine to coarse is initiated. The expansion process relaxes the matching conditions layer by layer in the following order: First-level expansion: Only the highest-level code c_0 is required to be the same, i.e., candidate queries q satisfy c_0(q) = c_0(p), without restrictions on lower-level codes c_1, ..., c_{m-1}. This operation expands the retrieval scope from exact matches to all queries under the same coarse-grained semantic category. For example, queries with the same technology but different lower-level codes, such as "smart home renovation" and "smart TV recommendation", are included.

[0062] Second-level expansion: If the number of candidates is still insufficient after the first-level expansion, then the codes c_0 and c_1 of the first two levels are further required to be identical. That is, the candidate query q must satisfy c_0(q) = c_0(p) and c_1(q) = c_1(p). There are no restrictions on c_2 and lower levels. This operation narrows the search scope to queries within the same medium-granularity category.

[0063] The expansion process continues until the number of candidates reaches a preset upper limit or has been expanded to the coarsest level. The query set obtained from each level of expansion is denoted as S_coarse_i. Finally, the results of all levels of expansion are merged into S_coarse, which is the second candidate query result mentioned above.

[0064] Step S2034: Merge and deduplicate the first candidate query result and the second candidate query result to obtain the corresponding candidate query set.

[0065] In this embodiment, a preliminary candidate pool S_raw = S_fine ∪ S_coarse is obtained by performing a union operation between the exact matching set S_fine and the coarse-grained extended set S_coarse. Then, deduplication is performed on S_raw: if two queries have completely identical text, one is retained; if the texts are different but semantically highly similar (which can be determined by the cosine similarity of the reconstructed vectors using RQ-VAE, and considered synonymous if the similarity exceeds a threshold), the query with the higher historical click-through rate is retained. After deduplication, the remaining candidates are initially sorted from highest to lowest historical click-through rate. The click-through rate is calculated as CTR(q) = N_click(q) / N_impression(q), where N_click(q) is the historical click count of query q, and N_impression(q) is its historical impression count. Click-through rate reflects the popularity of the query among real users and can be used as a proxy indicator of relevance.

[0066] Finally, the top-K items are truncated from the sorted candidate list to form the final candidate set H_p (i.e., the candidate query set), where the value of K is typically set to 3 to 5. For example, the truncated result might be H_p = ["smartphone latest price", "smart watch battery life test", "smart home renovation"]. This candidate set will be used as an enhanced reference sequence input into the next step of the multi-source information fusion encoding process.

[0067] Step S2035: The candidate query set is used as the candidate query list.

[0068] Based on the above processing flow, the essence of the RQ-VAE-based candidate set retrieval method approved in this application is to transform semantic retrieval from approximate nearest neighbor search in a continuous vector space to precise hash lookup and hierarchical expansion in a discrete code space. Precise matching utilizes complete semantic IDs to achieve constant-level retrieval, ensuring high relevance of the recalled results; coarse-grained expansion supplements candidate diversity by relaxing matching conditions layer by layer, while maintaining controllable semantic drift. This two-stage retrieval combination enables the system to obtain sufficient and high-quality candidate queries within millisecond latency, providing accurate contextual references for subsequent autoregressive generation processing, while avoiding the semantic drift problem caused by approximate calculations in traditional vector retrieval.

[0069] In some alternative implementations, step S205 includes the following steps: Step S2051: Based on a preset specified order, the prefix data, the candidate query list, the historical query sequence, and the portrait data are concatenated to obtain the corresponding input sequence.

[0070] In this embodiment, a unified input sequence is formed by concatenating four types of heterogeneous information in a specified order and inserting separator markers. The specific organization is as follows: x_u = {t_cls, p, t_sep, H_p, t_sep, H_u, t_sep, U}. Here, p is the prefix data, H_p is the retrieved candidate query list (each candidate is arranged as a text fragment in sequence), H_u is the user's historical query sequence (arranged in reverse chronological order, with the most recent query first), U is the textual description of the user profile, i.e., profile data (e.g., converting the structured tag "technology enthusiast, interested in consumer electronics" into a natural language description), t_cls is the classification start marker, and t_sep is the separator marker used to identify the boundaries of different data.

[0071] The specified order is not arbitrary but designed: the prefix p is located closest to t_cls to ensure the highest global visibility in the self-attention mechanism; the retrieval candidate H_p follows immediately to ensure that its supplementary role to the prefix is ​​captured early in the encoding process; the historical query H_u is located after the candidate, and the portrait U is located at the very end. This arrangement from near to far and from concrete to abstract conforms to the Transformer's utilization of positional information.

[0072] Step S2052: Perform word embedding mapping and position encoding superposition processing on the input sequence to obtain the corresponding embedding sequence.

[0073] In this embodiment, the specific implementation process of performing word embedding mapping and position encoding superposition processing on the input sequence to obtain the corresponding embedded sequence will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0074] Step S2053: Invoke the pre-built encoder.

[0075] In this embodiment, the encoder consists of L layers (usually L is 6 to 12 layers) of stacked Transformer blocks. Each layer contains a multi-head self-attention sublayer and a feedforward neural network sublayer. The two sublayers contain residual connections and layer normalization operations.

[0076] Step S2054: Based on a preset self-attention mechanism, the encoder is used to perform information fusion processing on the embedded sequence to obtain the corresponding fused data.

[0077] In this embodiment, the implementation process of using the encoder to perform information fusion processing on the embedded sequence based on a preset self-attention mechanism includes: 1) Multi-head Self-Attention Sublayer: For the l-th layer (1 ≤ l ≤ L), the output H^(l-1) of the previous layer is first mapped to the query matrix Q, key matrix K, and value matrix V through three learnable linear projection matrices W_Q, W_K, and W_V, respectively. Then, the attention weights are calculated using the formula Attention(Q, K, V) = softmax(QKT / √d_k) × V, where d_k is the dimension of the key vector. The softmax function normalizes along the key vector dimension, ensuring that the sum of the attention weights for each query term to all key terms is 1. The multi-head mechanism splits Q, K, and V into h heads (usually h = 8 or 12). Each head independently calculates attention in a d_k / h dimensional subspace. The outputs of each head are then concatenated and linearly projected back to d dimensions. This mechanism allows the model to capture multiple relationships between terms simultaneously from different representation subspaces.

[0078] 2) Information Interaction Process: Under the self-attention mechanism, any two terms in the sequence can be directly connected, regardless of distance. For example, a term in the prefix p at the beginning of the sequence can directly focus on terms in the historical query H_u at the end of the sequence, thus incorporating the user's past search intent into the understanding of the current prefix. Similarly, descriptive terms in the user profile U can also directly influence the representation of the prefix p. After L layers of encoding, the representation of each input term has fully absorbed the global context information.

[0079] 3) Prefix Representation Extraction: From the output H^L of the final layer of the encoder, extract the hidden state vector corresponding to the position of the prefix p, denoted as h_ctx ∈ R^d. This vector is the context representation of the entire input sequence after deep fusion, condensing the prefix semantics, supplementary information of retrieval candidates, user historical behavior, and the entire content of the user profile.

[0080] Step S2055: Use the fused data as the context representation.

[0081] In this embodiment, during the construction of the aforementioned embedded sequence, each candidate query in the candidate query list H_p is encoded using RQ-VAE, and its hierarchical semantic ID has a strict correspondence with the semantic ID of the prefix p in the hierarchical structure. When the encoder performs self-attention fusion on the complete sequence containing H_p, the lexical representations of each candidate query in H_p are directly applied to the lexical representation of the prefix p through attention weights. Since the semantic ID of H_p is consistent with the semantic ID of the prefix at a high level (as guaranteed by the fine-to-coarse retrieval strategy), this attention interaction naturally injects the hierarchical semantic structure information provided by RQ-VAE into h_ctx. Specifically, the encoder learns during the fusion process that the representation of the prefix p should be close to the representations of the candidate queries in H_p in the semantic space, while remaining far away from the representations of irrelevant queries. This constraint is equivalent to providing a "semantic anchor" for the model, ensuring that the generated result does not deviate from the core semantic direction of the prefix, and maintaining semantic consistency even under the influence of user history or profile information.

[0082] Based on the above processing flow, the essence of the multi-source information fusion encoding processing provided in this application is to complete the deep interaction and unified representation of heterogeneous information in a single forward computation. In traditional multi-stage cascaded architectures, the recall module only processes the matching of prefixes and candidates, and the ranking module only processes the matching of candidates and user features. Information is fragmented between stages, making it impossible to achieve cross-stage semantic collaboration. This application, through a unified encoder, places four types of information—prefixes, subsequent candidates, user historical queries, and user profiles—in the same sequence, and utilizes a multi-head self-attention mechanism to enable direct interaction between any two types of information. This fully connected information fusion method allows specific behaviors in the user's history to directly correct the understanding of the current prefix, long-term preferences in the user profile to directly influence the generation direction, and RQ-VAE retrieval candidates to ensure that the generated results do not deviate from the core semantics of the prefix through a semantic anchoring mechanism. The final output context representation is a unified semantic representation that integrates all information sources, which is beneficial for providing complete and consistent conditional inputs for the autoregressive generation of the decoder.

[0083] In some alternative implementations, step S2052 includes the following steps: Step S20521: Call the preset lexical embedding lookup table.

[0084] In this embodiment, the lexical embedding lookup table is a pre-constructed data table that stores the mapping relationship between lexical units and continuous vectors.

[0085] Step S20522: Based on the lexical embedding lookup table, map the lexical units in the input sequence into continuous vectors of fixed dimensions to obtain the corresponding lexical embedding matrix.

[0086] In this embodiment, for each word in the constructed input sequence x_u, it is mapped to a continuous vector of fixed dimension d through the word embedding lookup table mentioned above, so as to obtain the word embedding matrix E ∈ R^(|x_u|)×d, where |x_u| is the total length of the sequence.

[0087] Step S20523: Obtain the preset position encoding matrix.

[0088] In this embodiment, the aforementioned position encoding matrix refers to a learnable position encoding matrix P ∈ R^(|x_u|)×d.

[0089] Step S20524: Based on the position encoding matrix, the word embedding matrix is ​​superimposed to obtain the corresponding initial representation matrix.

[0090] In this embodiment, since the Transformer architecture itself does not contain prior knowledge of sequence order, a learnable positional encoding matrix P ∈ R(|x_u|)×d needs to be superimposed on the embedding matrix. Each position i in the positional encoding corresponds to a unique vector p_i, which is generated by different frequency combinations of sine and cosine functions, enabling the model to perceive the relative distance between any two words. After superposition, an initial representation matrix H_0 = E + P ∈ R(|x_u|)×d is obtained. This initial representation matrix carries both semantic and positional information of each word and serves as the input to the encoder (i.e., the embedding sequence).

[0091] Step S20525: Use the initial representation matrix as the embedding sequence.

[0092] Based on the above processing flow, this application uses a lexical embedding lookup table to map the lexical units in the input sequence into a continuous vector of fixed dimensions to obtain a lexical embedding matrix. Then, it uses a positional encoding matrix to superimpose the lexical embedding matrix and uses the resulting initial representation matrix as the corresponding embedding sequence. This enables the automatic and accurate completion of lexical embedding mapping and positional encoding superposition processing for the input sequence, effectively ensuring the accuracy and standardization of the generated embedding sequence, and helping to improve the compliance of data input for the encoder.

[0093] In some optional implementations of this embodiment, step S206 includes the following steps: Step S2061: Invoke the pre-built decoder.

[0094] In this embodiment, the above-mentioned decoder is also composed of L layers of stacked Transformer blocks, where the number of layers is usually the same as that of the encoder. Each layer comprises three sub-layers: a causal mask-aided self-attention sub-layer, a cross-attention sub-layer, and a feed-forward neural network sub-layer. Among them, the causal mask-aided self-attention sub-layer ensures that when generating the token at step t, the decoder can only attend to the previous t-1 generated tokens and cannot access subsequent tokens. This is achieved by setting the attention scores of future positions to negative infinity during attention weight calculation. The cross-attention sub-layer allows the decoder to directly attend to the entire output sequence of the encoder during each generation step, thereby injecting the multi-source context information fused by the encoder into the current generation step. Specifically, the query for cross-attention comes from the hidden state of the decoder itself, while the keys and values come from the encoder output H^L, which enables the decoder to dynamically extract the information most relevant to the current generation from the global context provided by the encoder when generating each token.

[0095] Step S2062: based on the decoder, a preset beam search strategy is used to perform autoregressive stepwise generation processing on the context representation to obtain corresponding complete query sequences, wherein the number of the complete query sequences is plural.

[0096] In this embodiment, the beam width B of the above-mentioned beam search strategy is set to a preset value (usually B = 5 to 10), which means that the system maintains B independent generation paths simultaneously. At the initial moment, all B paths start with the start token t_start as the first token, and each path has an independent hidden state sequence and cumulative probability record.

[0097] wherein, the implementation process of the above autoregressive stepwise generation processing includes: Step 1: receiving the context representation h_ctx ∈ R^d output by the final layer of the encoder as an initial condition.

[0098] Step 2: calculating the probability distribution at step t. In the t-th step of autoregressive generation, for each path k (k = 1, 2, ..., B) in beam search, the decoder calculates the conditional probability distribution of each token in the vocabulary V as the next token based on the partial sequence y_{<t}^{(k)} already generated by the path and the context representation h_ctx of the encoder.

[0099] First, the generated sequence y_{<t}{(k)} is fed into a decoder, and after layer-by-layer calculation through L layers, the decoder hidden state h_t{(k)} ∈ Rd at the current step is obtained. Subsequently, h_t{(k)} is mapped to a logit score vector of vocabulary dimension through a linear projection layer, and then normalized into a probability distribution through a softmax function. The probability calculation formula is P(y_t= v | y_{<t}{(k)}, h_ctx) = softmax(W_o × h_t{(k)} + b_o)[v], where W_o ∈ R{|V|×d} is a learnable output projection matrix, b_o ∈ R{|V|} is a learnable bias vector, and v is any token in the vocabulary. This distribution represents the probability that each token becomes the next token under the condition of the generated content of the current path and the global context.

[0100] Step 3: Path expansion and pruning for beam search. For each path k, select the Top-B tokens with the highest probabilities from its probability distribution as candidate expanded tokens. B paths each generate B candidate tokens, resulting in a total of B× B candidate expansions. For each candidate expansion, calculate its cumulative log probability of the path. The cumulative log probability of path k after selecting token v at step t is updated as: log P^{(k)}(y_{≤t}) = log P^{(k)}(y_{<t}) +log P(y_t = v | y_{<t}^{(k)}, h_ctx), where log P^{(k)}(y_{<t}) is the cumulative log probability of the first t-1 steps of the path, and log P(y_t = v | y_{<t}^{(k)}, h_ctx) is the log probability of selecting token v at the current step. Furthermore, among all B × B candidate expansions, B paths with the highest cumulative log probabilities are selected as surviving paths to enter step t+1, and the remaining paths are pruned and discarded. This selection process ensures that B globally optimal generation trajectories are retained at each step, rather than making greedy choices only locally.

[0101] By repeating the above expansion and pruning process until all surviving paths have generated an end token t_end, or a preset maximum generation length T_max is reached. If a path generates t_end in advance, the path is terminated and no longer participates in subsequent expansions.

[0102] Step S2063: Obtain the cumulative log probability of each complete query sequence, and perform sorting processing on all the complete query sequences based on the cumulative log probability to obtain a corresponding query sequence list.

[0103] In this embodiment, at most B complete query sequences are obtained after the beam search. Since some paths may terminate prematurely due to the generation of `t_end`, the actual number of complete output sequences may be less than B. A list of query sequences is obtained by sorting all complete query sequences from highest to lowest cumulative log probability. A higher cumulative log probability indicates a greater overall generation probability of the sequence under given context conditions, meaning a higher confidence level of the model in the result.

[0104] Step S2064: Select a specified number of query sequences with the highest cumulative log probability from the query sequence list.

[0105] In this embodiment, the top-N sequences with the highest cumulative log probability are selected from the sorted query sequence list as the final output, where N is typically set to 5 to 10. For example, the sorted results (i.e., the query sequence list) might be: "Latest smartphone prices" (cumulative log probability -12.3), "Smartphone recommendations for 2026" (cumulative log probability -13.1), "Smart watch and phone comparison" (cumulative log probability -14.5), "Smart home furnishing brand ranking" (cumulative log probability -15.0), and "Smart investment and financial management basics" (cumulative log probability -16.2). These sequences will be presented as query completion suggestions recommended to the user by the system, arranged in descending order of confidence.

[0106] Step S2065: Integrate all the specified query sequences to obtain the corresponding first generated list.

[0107] In this embodiment, all specified query sequences can be stored sequentially into a blank list in the order of selection, and the resulting first generated list can be used as the corresponding query list.

[0108] Step S2066: Use the first generated list as the query list.

[0109] Based on the above processing flow, this application provides an autoregressive generation process that transforms the unified semantic representation output by the encoder into specific natural language text. Its core mechanism is as follows: In each generation step, the decoder extracts the most relevant information from the encoder's global context through cross-attention, while ensuring left-to-right consistency in the generation process through causal self-attention. The bundle search strategy, by simultaneously maintaining multiple generation paths and performing global optimal pruning at each step, avoids the global suboptimal problem caused by early local optimal choices in greedy search. The cumulative log probability is used as the ranking criterion to ensure that the output result statistically best matches the model's understanding of the given context. The entire generation process is completed in a single decoder forward computation. Combined with the parallel path maintenance of the bundle search, the total latency is controlled at the millisecond level, effectively meeting the real-time requirements of interactive search and thus improving the processing efficiency of data queries.

[0110] In some optional implementations of this embodiment, step S207 includes the following steps: Step S2071: Based on a preset multi-dimensional deduplication strategy, the query list is deduplicated to obtain the corresponding first query list.

[0111] In this embodiment, the multi-dimensional deduplication strategy refers to performing two layers of deduplication operations on the query list, including: precise deduplication: traversing all query sequences in the query list and performing a word-by-word comparison at the lexical level for each pair of queries. If the lexical sequences of two queries are completely identical, their cumulative log probabilities are compared, and the one with the higher cumulative log probability is retained, while the other is discarded. This operation ensures that there are no completely duplicated entries in the output list.

[0112] Synonymous Deduplication: For query pairs that were not precisely filtered out, their semantic similarity is further calculated. Specifically, each query is mapped to a fixed-dimensional vector representation using a lexical embedding model, and the cosine similarity between two query vectors is calculated. If the similarity exceeds a preset threshold (usually set to 0.85 to 0.92), the two queries are considered highly semantically similar. In this case, the query with the higher cumulative logarithmic probability is retained, and the other is discarded. For example, although "smartphone price" and "smartphone quote" have different lexical units, their semantics are highly overlapping; after synonym deduplication, only the query with the higher probability is retained.

[0113] After deduplication, a concise query list with no duplicates and no semantic redundancy is obtained.

[0114] Step S2072: Perform compliance filtering on the first query list to obtain the corresponding second query list.

[0115] In this embodiment, the process of performing compliance filtering on each query in the deduplicated first query list includes: 1) Violation word matching: The query text is matched against a preset violation word library and a business blacklist. The violation word library contains general risk terms such as sensitive words, prohibited words, and politically sensitive words; the business blacklist contains non-compliant expressions specific to the financial industry, such as "insider information" and "guaranteed returns". The matching method combines exact matching and keyword matching: First, it checks whether the query completely contains any entry in the blacklist. If it does, the query is directly removed. Second, it checks whether it contains any keyword in the blacklist. If it does and the keyword appears in the core semantic position (judged by the position weight of the word element in the sequence), it is also removed.

[0116] 2) Syntax Standards Check: For queries that pass the violation filtering, their compliance with basic syntactic standards is checked. Specifically, a lightweight language model is used to score the query's syntactic accuracy. If the score is below a preset threshold (indicating that the sequence is grammatically incorrect or does not conform to natural language conventions), the query is marked as invalid and discarded. Simultaneously, the query length is checked to ensure it is within a reasonable range (e.g., no more than 50 tokens). Queries that are too short or too long are filtered to ensure the readability of the displayed results.

[0117] After filtering, a list of compliant and grammatically correct candidate queries is obtained.

[0118] Step S2073: The second query list is reordered by confidence to obtain the corresponding third query list.

[0119] In this embodiment, the filtered candidate query list is finally sorted according to a comprehensive ranking score. The primary ranking criterion is the cumulative logarithmic probability log P(q) of each query. A higher cumulative logarithmic probability indicates a greater overall generation probability of the query under given context conditions, a higher confidence level of the model in the result, and a higher ranking. Then, based on the primary ranking criterion, a business rule score R_business(q) is introduced as an adjustment factor. R_business(q) is calculated as follows: keyword detection is performed on the query text. If it contains keywords of query types that users frequently click in financial scenarios (such as "price," "yield," "ranking," "rating," "ranking," etc.), a fixed positive score is given; if it contains low-frequency or low-conversion keywords, no score is given or a slight negative adjustment is given. This factor reflects the business side's preference for different query types.

[0120] Furthermore, the final ranking score for each query is calculated using the formula Score(q) = λ × log P(q) + (1 -λ) × R_business(q), where λ is a balancing coefficient, typically ranging from 0.7 to 0.9. A larger λ value indicates a higher weighting of model probabilities, resulting in a ranking result closer to the model's own judgment; a smaller λ value indicates a higher weighting of business rules, resulting in a ranking result closer to the business side's preferences. All candidate queries are ranked from highest to lowest based on their Score(q).

[0121] Step S2074: Extract the target query sequence of the target quantity from the third query list in sequence.

[0122] In this embodiment, the top-N query sequences are extracted from the sorted list as the final output (i.e., the target query sequence), where N is typically set to 5 to 10. The extraction operation directly takes the first N items from the sorted list without further filtering.

[0123] Step S2075: Integrate the target query sequence to obtain the corresponding second generated list.

[0124] In this embodiment, all target query sequences can be stored sequentially into a blank list according to the order of extraction, and the resulting second generated list can be used as the corresponding target query list.

[0125] Step S2076: Use the second generated list as the target query list.

[0126] Based on the above processing flow, the post-processing and sorting processes provided in this application are the final quality control steps before generating query results for users. Deduplication eliminates redundant entries in the query results, ensuring the simplicity of the displayed list; compliance filtering constitutes the last line of defense for content security, ensuring that all queries presented to users comply with regulatory requirements and language standards; confidence-based re-sorting establishes an adjustable balance mechanism between the model's own probability judgment and the operational needs of the business side, ensuring that the final presented results have both statistically high confidence and conform to business preferences in financial scenarios. The computational overhead of the entire post-processing process is extremely low and will not have a substantial impact on the system's real-time performance. This is thanks to the design advantage of the end-to-end generative architecture, which compresses the complex retrieval-sorting-generation process into a single forward computation.

[0127] In some optional implementations, the model training process of the generative model in this application includes: 1. The training implementation process of the RQ-VAE semantic ID generator (i.e., the RQ-VAE encoder or semantic generator) includes: Data Collection and Pairing: A large number of "user input prefix - final click query" pairs are extracted from the historical interaction logs of the search system. For example, the prefix "smart" is ultimately clicked by the user in the complete query "smartphone latest price". For each pair, the prefix text and the complete query text are converted into token sequences using a tokenizer. Then, a pre-trained language model (such as BERT) encodes the token sequences into fixed-dimensional continuous vectors, resulting in a prefix embedding vector e_p and a query embedding vector e_q. The cosine similarity between the two is calculated as sim(e_p, e_q) = (e_p · e_q) / (‖e_p‖ × ‖e_q‖). Only pairs with a similarity exceeding a preset threshold τ (usually τ is set to 0.75 to 0.85) are retained, filtering out semantically irrelevant noise data to ensure the quality of the training data. Simultaneously, query embeddings are deduplicated to prevent the model from biasing towards high-frequency queries due to repeated occurrences of the same query.

[0128] RQ-VAE Model Training: The RQ-VAE model is trained using a residual quantization variational autoencoder on high-quality query embeddings. The encoder receives the input vector x and encodes it progressively through multiple residual quantization modules. In the i-th layer (i ranges from 0 to m-1, where m is the number of quantization layers, typically 3 to 5), the encoder outputs a continuous vector r_i. Then, it searches the codebook of that layer for the centroid vector e_{c_i} that has the closest Euclidean distance to r_i, calculates the residual r_i - e_{c_i}, and uses this residual as the input for the next layer. Finally, the input vector x is represented as a combination of m discrete codes {c_0, c_1, ..., c_{m-1}}. The decoder receives this set of discrete codes and reconstructs the original vector x by upsampling layer by layer. During training, two loss functions are optimized simultaneously. The first is the reconstruction loss: L_recon = ||x - x|| The first loss is ||², which ensures the decoder can accurately recover the original vector from the discrete code. The second is the quantization loss: L_rqvae = ∑(i=0 to m-1) [||sg[r_i] - e_{c_i}||² + β||r_i - sg[e_{c_i}]||²], where sg[·] represents the stop gradient operation, and β is the balance coefficient (usually 0.25). This loss ensures that the quantized code of each layer is as close as possible to the corresponding centroid, while minimizing the residual. The total loss is L(x) = L_recon + L_rqvae. Training uses the Adam optimizer, with the learning rate starting at 1e-4 and gradually decreasing using a cosine annealing strategy. The number of training epochs is usually 50 to 100 until the loss converges.

[0129] Semantic codebook (i.e., offline index) construction: After training, all centroid vectors {e_{c_i}} in each layer of the RQ-VAE codebook are extracted and organized hierarchically into a layered semantic codebook. The centroids of layer 0 represent the coarsest-grained semantic categories (e.g., "technology", "finance", "consumer"), layer 1 represents medium-grained categories (e.g., "electronic products", "financial tools"), and layer m-1 represents the finest-grained categories (e.g., "smartphones", "money market funds"). This codebook is then frozen and no longer participates in subsequent training, serving as an offline index for the inference phase.

[0130] 2. The training process of the generative master model M_t includes: Training Sample Construction: For each historical interaction record, a training sample is constructed. The sample contains four parts of input. The first part is the current prefix p, which is the character sequence typed by the user in the input box, such as "smart". The second part is the enhanced reference sequence H_p, which uses a pre-trained RQ-VAE to encode the prefix to obtain its semantic ID. Then, a fine-to-coarse strategy is used to retrieve the top-K semantically related historical queries (e.g., K=5) from the semantic codebook. These queries are sorted in descending order of relevance as H_p. The third part is the user's historical query sequence H_u, which takes the user's most recent N search records (e.g., N=10) and sorts them in reverse chronological order, with the most recent one at the top. The fourth part is the user profile U, which converts the user's structured tags (e.g., "tech enthusiast", "interested in consumer electronics", "financial management needs") into natural language descriptive text through a template. The target output Q is the complete query that the user finally clicked. Organize all inputs into a unified sequence according to a special label: x_u = {t_cls, p, t_sep, H_p, t_sep, H_u, t_sep, U}, where t_cls is the sequence start label and t_sep is the separator label.

[0131] End-to-end training: The model adopts a Transformer encoder-decoder architecture. The encoder receives the aforementioned input sequence, first obtains token embedding vectors by performing embedding lookup for each token, then adds learnable position encodings. The encoder is composed of L layers (typically L ranges from 6 to 12) of stacked multi-head self-attention layers and feed-forward neural network layers. In each layer, connections between any two tokens are established through the self-attention mechanism, and the calculation formula for attention weights is Attention(Q, K, V) =softmax(QK^T / √d_k) × V, where Q = W_q × h, K = W_k × h, V = W_v × h, and d_k is the dimension of the key vector. After L layers of encoding, each input token carries global context information. The decoder is also composed of L layers, each of which includes a masked self-attention layer (to prevent seeing future tokens) and a cross-attention layer (to focus on the output of the encoder). The decoder generates target queries in an autoregressive manner. At the t-th step, based on the previous t-1 generated tokens and the context representation h_ctx of the encoder, the probability distribution of the next token is calculated as: P(y_t | y_{<t}, h_ctx) = softmax(W_o× h_t^dec + b_o), where h_t^dec is the hidden state of the decoder at the t-th step, and W_o and b_o are learnable parameters. Cross-entropy loss is used during training: L_ce = -∑(t=1 to T) log P(y_t^* | y_{<t}^*, h_ctx), where y_t^* is the ground-truth target token, and T is the length of the target sequence. In addition, an auxiliary semantic anchor loss is introduced: the prefix representation h_p^enc output by the encoder is aligned with the semantic ID obtained by encoding the prefix via RQ-VAE, and the contrastive loss is calculated as L_anchor = -log [exp(sim(h_p^enc, e_id) / τ) / ∑(j) exp(sim(h_p^enc, e_j) / τ)], where e_id is the centroid of the correct semantic ID, e_j is other centroids in the codebook, and τ is the temperature coefficient. The total loss is L_total = L_ce + λ_anchor × L_anchor, and λ_anchor is typically set between 0.1 and 0.3. The AdamW optimizer is adopted for training, and a warmup strategy is used for the learning rate: the learning rate linearly rises to the peak (e.g., 5e-5) in the first 10% of steps, then decays via cosine annealing, and the number of training epochs ranges from 10 to 30.

[0132] Seed model M_t acquisition: When the total loss converges and the generation quality metrics (such as BLEU, ROUGE) on the validation set no longer improve, the model parameters are saved to obtain the seed model M_t. At this point, M_t has the ability to generate complete queries based on prefixes and multi-source context, but its generation results have not yet been jointly optimized for their dependence on the candidates retrieved by RQ-VAE.

[0133] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0134] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0135] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0136] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of an artificial intelligence-based data query device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0137] like Figure 3 As shown, the artificial intelligence-based data query device 300 described in this embodiment includes: a receiving module 301, an extraction module 302, a retrieval module 303, an acquisition module 304, a fusion module 305, a generation module 306, and a processing module 307. Wherein: The receiving module is used to receive prefix data entered by the user in the search box; The extraction module is used to perform semantic extraction processing on the prefix data based on a preset semantic generator to obtain the corresponding semantic data; The retrieval module is used to perform retrieval processing on the semantic data based on a preset clustering search strategy to obtain the corresponding candidate query list; The acquisition module is used to acquire the historical query sequence corresponding to the user, and to acquire the user's profile data; The fusion module is used to perform fusion processing on the prefix data, the candidate query list, the historical query sequence and the profile data based on a preset encoder to obtain the corresponding context representation; The generation module is used to perform autoregressive stepwise generation processing on the context representation based on a preset decoder to obtain the corresponding query list; The processing module is used to optimize the query list based on a preset post-processing strategy to obtain the corresponding target query list, and output the target query list.

[0138] In some optional implementations of this embodiment, the extraction module 302 includes: The word segmentation submodule is used to segment the prefix data to obtain the corresponding word sequence. The first processing submodule is used to obtain a preset specified tag and insert the specified tag into a specified position in the lexical sequence to obtain the corresponding lexical embedding sequence; The first calling submodule is used to invoke the pre-built semantic generator; The computational submodule is used to perform forward computation processing on the word embedding sequence based on the semantic generator to obtain the corresponding discrete semantic representation; The first determining submodule is used to use the discrete semantic representation as the semantic data.

[0139] In some optional implementations of this embodiment, the retrieval module 303 includes: The second submodule is used to invoke a pre-built offline index; The first retrieval submodule is used to perform precise matching and retrieval of the semantic data based on the offline index to obtain the corresponding first candidate query result; The second retrieval submodule is used to perform coarse-grained extended retrieval of the semantic data based on the offline index to obtain the corresponding second candidate query results; The second processing submodule is used to merge and deduplicate the first candidate query result and the second candidate query result to obtain the corresponding candidate query set; The second determining submodule is used to use the candidate query set as the candidate query list.

[0140] In some optional implementations of this embodiment, the fusion module 305 includes: The splicing submodule is used to splice the prefix data, the candidate query list, the historical query sequence, and the profile data according to a preset specified order to obtain the corresponding input sequence; The third processing submodule is used to perform word embedding mapping and position encoding superposition processing on the input sequence to obtain the corresponding embedding sequence; The third submodule is used to invoke the pre-built encoder; The fusion submodule is used to perform information fusion processing on the embedded sequence using the encoder based on a preset self-attention mechanism to obtain corresponding fused data; The third determining submodule is used to use the fused data as the context representation.

[0141] In some optional implementations of this embodiment, the third processing submodule includes: The calling unit is used to call the preset lexical embedding lookup table; The mapping unit is used to map the lexical units in the input sequence into continuous vectors of fixed dimensions based on the lexical embedding lookup table, so as to obtain the corresponding lexical embedding matrix. The acquisition unit is used to acquire a preset position encoding matrix; The processing unit is used to perform superposition processing on the word embedding matrix based on the position encoding matrix to obtain the corresponding initial representation matrix; A determining unit is used to take the initial representation matrix as the embedding sequence.

[0142] In some optional implementations of this embodiment, the generation module 306 includes: The fourth submodule is used to invoke a pre-built decoder; The fourth processing submodule is used to perform autoregressive stepwise generation processing on the context representation based on the decoder using a preset beam search strategy to obtain the corresponding complete query sequence; wherein, the number of complete query sequences is multiple; The first sorting submodule is used to obtain the cumulative log probability of each complete query sequence, and sort all the complete query sequences based on the cumulative log probability to obtain the corresponding query sequence list. The first filtering submodule is used to filter a specified number of specified query sequences with the highest cumulative log probability from the query sequence list; The first integration submodule is used to integrate all the specified query sequences to obtain the corresponding first generated list; The fourth determining submodule is used to use the first generated list as the query list.

[0143] In some optional implementations of this embodiment, the processing module 307 includes: The deduplication submodule is used to perform deduplication processing on the query list based on a preset multi-dimensional deduplication strategy to obtain the corresponding first query list. The filtering submodule is used to perform compliance filtering on the first query list to obtain the corresponding second query list; The second sorting submodule is used to perform confidence reordering on the second query list to obtain the corresponding third query list; The second filtering submodule is used to extract the target query sequence of the target quantity sequentially from the third query list; The second integration submodule is used to integrate the target query sequence to obtain the corresponding second generated list; The fifth determining submodule is used to use the second generated list as the target query list.

[0144] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0145] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0146] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0147] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for data querying methods based on artificial intelligence. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.

[0148] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions for the artificial intelligence-based data query method.

[0149] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.

[0150] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the artificial intelligence-based data query method described above.

[0151] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0152] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

[0153] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.

Claims

1. A data query method based on artificial intelligence, characterized in that, Includes the following steps: Receive prefix data entered by the user in the search box; The prefix data is semantically extracted based on a preset semantic generator to obtain the corresponding semantic data. The semantic data is retrieved and processed based on a preset clustering search strategy to obtain a corresponding candidate query list; Obtain the historical query sequence corresponding to the user, and obtain the user's profile data; The prefix data, the candidate query list, the historical query sequence, and the profile data are fused based on a preset encoder to obtain the corresponding context representation; The context representation is generated stepwise by autoregression based on a preset decoder to obtain the corresponding query list; The query list is optimized based on a preset post-processing strategy to obtain the corresponding target query list, and the target query list is output.

2. The data query method based on artificial intelligence according to claim 1, characterized in that, The step of performing semantic extraction processing on the prefix data based on a preset semantic generator to obtain the corresponding semantic data specifically includes: The prefix data is segmented to obtain the corresponding word sequence; Obtain a preset designated marker and insert the designated marker into a designated position in the lexical sequence to obtain the corresponding lexical embedding sequence; Invoke the pre-built semantic generator; Based on the semantic generator, the word embedding sequence is processed by forward computation to obtain the corresponding discrete semantic representation; The discrete semantic representation is used as the semantic data.

3. The data query method based on artificial intelligence according to claim 1, characterized in that, The step of retrieving and processing the semantic data based on a preset clustering search strategy to obtain the corresponding candidate query list specifically includes: Invoke a pre-built offline index; Based on the offline index, the semantic data is precisely matched and retrieved to obtain the corresponding first candidate query result; Based on the offline index, coarse-grained extended retrieval of the semantic data is performed to obtain the corresponding second candidate query results; The first candidate query result and the second candidate query result are merged and deduplicated to obtain the corresponding candidate query set; The candidate query set is used as the candidate query list.

4. The data query method based on artificial intelligence according to claim 1, characterized in that, The step of fusing the prefix data, the candidate query list, the historical query sequence, and the profile data based on a preset encoder to obtain the corresponding context representation specifically includes: The prefix data, the candidate query list, the historical query sequence, and the profile data are concatenated according to a preset specified order to obtain the corresponding input sequence; The input sequence is subjected to word embedding mapping and position encoding superposition processing to obtain the corresponding embedding sequence; Invoke the pre-built encoder; Based on a preset self-attention mechanism, the encoder is used to perform information fusion processing on the embedded sequence to obtain corresponding fused data; The fused data is used as the context representation.

5. The data query method based on artificial intelligence according to claim 4, characterized in that, The step of performing word embedding mapping and positional encoding superposition processing on the input sequence to obtain the corresponding embedding sequence specifically includes: Call the preset lexical embedding lookup table; Based on the lexical embedding lookup table, the lexical units in the input sequence are mapped to continuous vectors of fixed dimensions to obtain the corresponding lexical embedding matrix; Obtain the preset position encoding matrix; The word embedding matrix is ​​superimposed based on the position encoding matrix to obtain the corresponding initial representation matrix; The initial representation matrix is ​​used as the embedding sequence.

6. The data query method based on artificial intelligence according to claim 1, characterized in that, The step of performing autoregressive stepwise generation processing on the context representation based on a preset decoder to obtain the corresponding query list specifically includes: Invoke the pre-built decoder; Based on the decoder, the context representation is subjected to autoregressive stepwise generation using a preset beam search strategy to obtain the corresponding complete query sequence; wherein, the number of complete query sequences is multiple. Obtain the cumulative log probability of each complete query sequence, and sort all the complete query sequences based on the cumulative log probability to obtain the corresponding query sequence list; Filter the list of query sequences to select a specified number of query sequences with the highest cumulative log probability; All the specified query sequences are integrated to obtain the corresponding first generated list; Use the first generated list as the query list.

7. The data query method based on artificial intelligence according to claim 1, characterized in that, The step of optimizing the query list based on a preset post-processing strategy to obtain the corresponding target query list specifically includes: The query list is deduplicated based on a preset multidimensional deduplication strategy to obtain the corresponding first query list. The first query list is subjected to compliance filtering to obtain the corresponding second query list; The second query list is reordered by confidence to obtain the corresponding third query list; Extract the target query sequence of the target quantity from the third query list in sequence; The target query sequence is integrated to obtain the corresponding second generated list; The second generated list is used as the target query list.

8. A data query device based on artificial intelligence, characterized in that, include: The receiving module is used to receive prefix data entered by the user in the search box; The extraction module is used to perform semantic extraction processing on the prefix data based on a preset semantic generator to obtain the corresponding semantic data; The retrieval module is used to perform retrieval processing on the semantic data based on a preset clustering search strategy to obtain the corresponding candidate query list; The acquisition module is used to acquire the historical query sequence corresponding to the user, and to acquire the user's profile data; The fusion module is used to perform fusion processing on the prefix data, the candidate query list, the historical query sequence and the profile data based on a preset encoder to obtain the corresponding context representation; The generation module is used to perform autoregressive stepwise generation processing on the context representation based on a preset decoder to obtain the corresponding query list; The processing module is used to optimize the query list based on a preset post-processing strategy to obtain the corresponding target query list, and output the target query list.

9. A computer device, characterized in that, The system includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the data query method based on artificial intelligence as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the data query method based on artificial intelligence as described in any one of claims 1 to 7.