Data retrieval method and device, equipment and medium

By using standard databases and semantic relationship graphs in the data retrieval system, combined with user history behavior and intent recognition, the problem of insufficient understanding of user intent in existing technologies is solved, thereby improving the accuracy of data retrieval and user experience.

CN121256103APending Publication Date: 2026-01-02INFORMATION & COMM BRANCH OF STATE GRID JIANGSU ELECTRIC POWER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511376512.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing data retrieval methods cannot accurately understand users' semantic intent, resulting in inaccurate search results, poor scalability, and a poor user experience.

Method used

By acquiring the query information input by the user, semantic matching is performed using a pre-established standard database. Combined with the semantic relationship graph and the user's historical behavior, the user's interest vector and intent recognition results are determined, thereby identifying the target retrieval result from the associated retrieval results.

Benefits of technology

It enables proactive understanding of user query intent, improves the accuracy of data retrieval and user experience, and expands the relevance of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121256103A_ABST
    Figure CN121256103A_ABST
Patent Text Reader

Abstract

The invention discloses a data retrieval method and device, equipment and a medium. The method comprises the following steps: acquiring current query information input by a user; performing semantic matching on the current query information and a pre-established standard database, and determining an initial retrieval result; based on the initial retrieval result and a preset semantic relation graph, determining an associated retrieval result; determining a user interest vector and an intention recognition result according to user historical behaviors and the current query information; and determining a target retrieval result from the associated retrieval result through the user interest vector and the intention recognition result. According to the technical scheme, the retrieval result can be expanded, the query intention of the user can be recognized, and the problems that in the technical scheme in the prior art, the accuracy of data retrieval is low, and the use experience of the user is poor are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer information processing technology, and in particular to a data retrieval method, apparatus, device, and medium. Background Technology

[0002] With the continuous advancement of the construction of new power systems and digital grids, power grid companies have accumulated a large amount of data assets from different sources and modes in various power systems. Accurately and quickly retrieving the information needed by users from these massive data assets is crucial for performing tasks such as intelligent dispatching, fault diagnosis, equipment operation and maintenance, and data sharing.

[0003] Existing data retrieval methods typically match corresponding data assets based on keywords entered by the user. This approach fails to provide a deep semantic understanding of the query information or to determine the user's true intent. Furthermore, the data assets obtained through keyword matching have poor scalability and are ill-suited for handling complex queries.

[0004] Therefore, existing technical solutions suffer from low accuracy in data retrieval and poor user experience. Summary of the Invention

[0005] This invention provides a data retrieval method, apparatus, device, and medium that can expand retrieval results and identify user query intent, thereby solving the problems of low accuracy in data retrieval and poor user experience in existing technical solutions.

[0006] According to a first aspect of the present invention, a data retrieval method is provided, characterized in that it includes:

[0007] Get the current query information entered by the user;

[0008] The current query information is semantically matched with a pre-established standard database to determine the initial search results;

[0009] Based on the initial search results and the preset semantic relationship graph, determine the associated search results;

[0010] The user's interest vector and intent recognition result are determined based on the user's historical behavior and the current query information.

[0011] The target search result is determined from the associated search results based on the user interest vector and intent recognition results.

[0012] According to a second aspect of the present invention, a data retrieval apparatus is provided, characterized in that it comprises:

[0013] The acquisition module is used to acquire the current query information input by the user;

[0014] The matching module is used to perform semantic matching between the current query information and a pre-established standard database to determine the initial search results;

[0015] The association determination module is used to determine the association search results based on the initial search results and the preset semantic relationship graph;

[0016] The intent recognition module is used to determine the user's interest vector and intent recognition result based on the user's historical behavior and the current query information;

[0017] A filtering model is used to determine target search results from the associated search results based on the user interest vector and intent recognition results.

[0018] According to a third aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0019] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data retrieval method according to any embodiment of the present invention.

[0020] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the data retrieval method according to any embodiment of the present invention.

[0021] According to a fifth aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the data retrieval method described in any embodiment of the present invention.

[0022] The technical solution of this invention obtains the current query information input by the user, performs semantic matching between the current query information and a pre-established standard database to determine the initial search results, and then determines related search results based on the initial search results and a preset semantic relationship graph. This solves the problem of isolated user query results and expands related search results. Then, based on the user's historical behavior and the current query information, a user interest vector and intent recognition result are determined. Using the user interest vector and intent recognition result, the target search result is determined from the related search results. This solves the problem that traditional retrieval systems cannot proactively understand the user's true needs, achieves the recognition of the user's query intent, and improves the accuracy of data retrieval and user experience.

[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart of a data retrieval method provided according to Embodiment 1 of the present invention;

[0026] Figure 2 This is a flowchart of a data retrieval method provided according to Embodiment 2 of the present invention;

[0027] Figure 3 This is a schematic diagram of the structure of a data retrieval device according to Embodiment 3 of the present invention;

[0028] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the data retrieval method of this invention. Detailed Implementation

[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0031] Example 1

[0032] Figure 1 This is a flowchart illustrating a data retrieval method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations involving the retrieval of large amounts of multi-source heterogeneous data assets. The method can be executed by a data retrieval device, which can be implemented in hardware and / or software and configured within a system or platform that manages data assets. Figure 1 As shown, the method includes:

[0033] S101. Obtain the current query information input by the user.

[0034] The current query information can be the natural query language entered by the user.

[0035] For example, a data retrieval device may pre-provide a query input interface or component for the user, enabling the user to retrieve the data assets they desire by inputting natural query language based on the query input interface or component.

[0036] S102. Perform semantic matching between the current query information and a pre-established standard database to determine the initial search results.

[0037] The standard database can include semantic information corresponding to each data asset. The initial search results can be an initial set of data assets based on the semantic retrieval of the current query information.

[0038] Data assets can be a large amount of data resources from different sources and modalities that an enterprise accumulates in various systems. For example, in the power sector, data assets can include telemetry data, GIS spatial data, video images, alarm logs, work orders, text archives, etc., covering all aspects of power grid operation and maintenance, dispatching, monitoring, and fault handling.

[0039] For example, after obtaining the natural query language input by the user, the natural query language can be semantically encoded to determine the semantic vector corresponding to the natural query language. Then, the similarity between the semantic vector corresponding to the natural query language and the semantic information corresponding to each data asset in the standard database is calculated to determine the semantic information with high similarity and the corresponding data assets, thereby retrieving the initial set of data assets corresponding to the natural query language.

[0040] It should be noted that because data assets originate from different sources, they are diverse in type and complex in format, exhibiting a multi-source heterogeneous characteristic. Therefore, before retrieving data assets, these multi-source heterogeneous data assets can be represented using semantic information and a standard database can be built to facilitate data retrieval.

[0041] Optionally, the method for establishing the standard database may include:

[0042] Acquire multi-source heterogeneous data asset sets;

[0043] Multimodal feature encoding is performed on the data asset set to obtain the corresponding semantic information dataset;

[0044] The standard database is constructed based on the semantic information dataset.

[0045] The data asset set can include multiple data assets with different sources, formats, and types. The semantic information dataset can include the semantic information corresponding to each data asset.

[0046] For example, in the power industry, the management system or platform for power data assets can call interfaces with various power systems to obtain data assets accumulated in various power systems, thus obtaining a multi-source heterogeneous data asset set.

[0047] Then, data assets from different systems and formats can be standardized to achieve format unification, providing a consistent format for subsequent data processing and retrieval. Next, these format-unified data assets are classified (e.g., into text, image, time-series, and structured data), and features are extracted and encoded for each type, transforming them into unified semantic information to obtain the corresponding semantic information dataset. For example, text data assets can be encoded using BERT; image data assets can have visual features extracted using the ResNet-50 model; time-series data can be extracted using the Transformer Encoder model; and structured data can be embedded using TabNet. Furthermore, to unify the dimensions, a projection head (FC+ReLU) can be used to map the data from each modality to a unified embedding space and perform normalization.

[0048] Finally, the corresponding semantic information dataset is stored in the database according to a preset structure to build a standard database.

[0049] Understandably, by extracting and encoding features from different types of data assets, multimodal feature fusion can be achieved, enabling data retrieval of different types of data assets and expanding the application scope of data retrieval.

[0050] S103. Based on the initial search results and the preset semantic relationship graph, determine the associated search results.

[0051] A semantic relationship graph can be established based on the power business chain, representing the relationships between different data assets. For example, a semantic relationship graph may include at least two nodes and at least one edge. Nodes can represent data assets (e.g., equipment documents, measurement points, work orders, alarms, images, etc.). Edges can represent the relationships between different nodes (e.g., relationships such as "belonging," "dependency," "trigger," "same station / same line," etc.). The associated retrieval results may include associated data assets related to each initial data asset in the initial detection results. It is understood that this invention can automatically construct a graph structure based on the power grid business logic relationships, enabling the retrieval results to have semantic interpretability related to upstream and downstream businesses, enhancing the practicality and interpretability of the retrieval.

[0052] For example, after determining the initial search results, for each initial data asset in the initial search results, the node corresponding to that initial data asset can be identified as the starting point in the semantic relationship graph, and other nodes in the semantic relationship graph can be identified as candidate nodes. Then, based on the edges in the semantic relationship graph, a set of candidate paths from the starting point to each candidate node can be determined. Next, the path score of each candidate path in the candidate path set is calculated, thereby filtering out the target candidate nodes and the data assets corresponding to the target candidate nodes from the candidate paths with high path scores based on a preset number. Finally, the initial data assets and the data assets corresponding to the target candidate nodes are used as the associated search results.

[0053] Based on this, it is understandable that in the process of determining the associated search results, the business paths and related relationships represented by the candidate paths with high path scores can also be output as recommendation explanations, thereby realizing chain-like extended recommendations with path explanations.

[0054] Optionally, determining the associated search results based on the initial search results and the preset semantic relationship graph may further include:

[0055] Based on the initial search results, the target node is determined in the semantic relationship graph;

[0056] Based on the target node and the edges connected to the target node, determine the associated nodes;

[0057] The data assets corresponding to the target node and the associated node are identified as the associated retrieval results.

[0058] For example, based on each initial data asset in the initial search results, the target node corresponding to the initial data asset can be determined in the semantic relationship graph. Then, based on the target node and the edge connected to the target node, the associated node with the target node can be determined. Finally, the data assets corresponding to the target node and the associated node can be determined as the associated search results.

[0059] S104. Determine the user interest vector and intent recognition result based on the user's historical behavior and the current query information.

[0060] Among these, user history behavior can be a sequence of user behavior prior to inputting the current query. User interest vector can be a vector representing the user's long-term, diverse interests and preferences. Intent recognition result can be the intent category and confidence level corresponding to the current query, representing the user's immediate purpose for the current query (i.e., the current query information).

[0061] For example, based on the user behavior sequence in the user's historical behavior, the behavioral information of each user behavior can be determined, and the interest tags corresponding to the user behavior can be identified. Then, each interest tag is vectorized to form a user interest vector. Here, the interest tags can be scenario tags of the power business. For example, in the power industry, the names of different power systems can be used as interest tags, thereby identifying whether the user's historical behavior shows an interest in retrieving data assets related to various power systems.

[0062] Simultaneously, the current query information can be processed through word segmentation and stop word removal to determine the keywords. Then, these keywords are matched against a predefined intent dictionary to determine the intent recognition result corresponding to the current query information. The intent dictionary can include multiple preset intent fields representing different intents. For example, an intent dictionary can be constructed based on tasks such as intelligent scheduling, fault diagnosis, equipment maintenance, and data sharing of retrieved data assets, thereby identifying the task intent of the data assets retrieved in the current query information.

[0063] S105. Determine the target retrieval result from the associated retrieval results using the user interest vector and intent recognition result.

[0064] Understandably, after determining the user interest vector that can represent the user's long-term interests and preferences, and the intent recognition result that can represent the user's immediate purpose in the current query, the features of each data asset in the associated search results can be matched based on the user interest vector and intent recognition result to determine the data asset with matching features as the target search result, so as to achieve the goal of filtering the target search results that meet the user's personalized needs in the associated search results.

[0065] The technical solution of this invention obtains the current query information input by the user, performs semantic matching between the current query information and a pre-established standard database to determine the initial search results, and then determines related search results based on the initial search results and a preset semantic relationship graph. This solves the problem of isolated user query results and expands related search results. Then, based on the user's historical behavior and the current query information, a user interest vector and intent recognition result are determined. Using the user interest vector and intent recognition result, the target search result is determined from the related search results. This solves the problem that traditional retrieval systems cannot proactively understand the user's true needs, achieves the recognition of the user's query intent, and improves the accuracy of data retrieval and user experience.

[0066] Based on the above embodiments, this embodiment also provides an optional embodiment. This optional embodiment can further optimize step S102 of the above embodiments, which involves semantically matching the current query information with a pre-established standard database to determine the initial search results. Specifically, it may include:

[0067] The current query information is semantically encoded using a preset semantic encoding model to obtain a semantic vector of the current query information.

[0068] Based on knowledge graph technology, semantic constraints are applied to the current query information to obtain a restricted semantic subspace of the current query information;

[0069] Based on the restricted semantic subspace and semantic vector of the current query information, semantic matching is performed in the pre-established standard database to determine the initial search results.

[0070] The semantic encoding model can be the BERT model. Intelligent graph technology can be used to limit the semantic scope of the current query information.

[0071] For example, by using a pre-defined semantic encoding model, the current query information can be processed with contextual dynamic word vectors to obtain encoded semantic vectors. Simultaneously, a power grid knowledge graph can be introduced, and the entire power grid knowledge graph can be embedded and learned using the TransE algorithm to obtain vector representations of all entities and relationships in the graph, forming a restricted semantic subspace for the current query information.

[0072] Furthermore, based on the restricted semantic subspace and semantic vector of the current query information, the target semantic vector of the user's current query information within the restricted semantic subspace can be determined. Then, the similarity of the target semantic vector of the current query information within the restricted semantic subspace with the semantic information in the standard database can be matched. From the standard database, highly similar semantic information and corresponding data assets can be determined, and the initial data asset set can be retrieved.

[0073] The advantage of this setup is that it can utilize the structured and explicit prior knowledge in the knowledge graph to eliminate ambiguity in the user's current query information, clarify its semantic boundaries, and thus map the semantics of the current query information into a more precise and narrower restricted semantic subspace, thereby improving the accuracy of initial data retrieval.

[0074] Optionally, the step of performing semantic matching in the pre-established standard database based on the restricted semantic subspace and semantic vector of the current query information to determine the initial search result may further include:

[0075] Based on the restricted semantic subspace of the current query information, a candidate set is recalled in the standard database;

[0076] The semantic vector of the current query information is compared with the candidate set to calculate the similarity, and the initial retrieval result is determined based on the calculation result.

[0077] It should be noted that after determining the restricted semantic subspace and semantic vector of the current query information, we can first use the restricted semantic subspace of the current query information to recall semantic information that matches the restricted semantic subspace in the standard database, forming a candidate set. Then, we can construct a similarity matrix based on the semantic vector of the current query information and the semantic information in the candidate set, calculate the cosine similarity between the semantic vector of the current query information and the semantic information in the candidate set, and thus select the data assets corresponding to the semantic information with high similarity based on a preset number, as the initial retrieval results.

[0078] The advantage of this setup is that it allows for two-stage matching of the standard database using the limited semantic subspace and semantic vector of the current query information, further improving the accuracy of initial data retrieval.

[0079] Example 2

[0080] Figure 2 This is a flowchart of a data retrieval method provided in Embodiment 2 of the present invention. This embodiment can further optimize step S103 in Embodiment 1 above, and may include: determining a user interest vector based on the user's historical behavior and a preset time decay weight; determining the intent recognition result of the current query information based on a preset intent recognition classifier and the user interest vector; the intent recognition result includes: intent category and confidence level. Figure 2 As shown, the method includes:

[0081] S201. Obtain the current query information input by the user.

[0082] S202. Perform semantic matching between the current query information and a pre-established standard database to determine the initial search results.

[0083] S203. Based on the initial search results and the preset semantic relationship graph, determine the associated search results.

[0084] S204. Determine the user interest vector based on the user's historical behavior and the preset time decay weight.

[0085] The time decay weight represents the importance of behavior at different times. For example, user historical behavior at the most recent time when the current query information was obtained is the most important, and its corresponding weight is higher. Conversely, the weight of user historical behavior further back in time from when the current query information was obtained is lower. User interest vectors can include labels related to power business scenarios, such as the name of the power system.

[0086] For example, based on user historical behavior, by determining the power business scenario tags that users are interested in (such as determining the data source and type of data assets that users are interested in), each user behavior in the user behavior sequence is mapped to a specific interest tag, thereby obtaining each historical interest vector corresponding to the user's historical behavior. Then, based on the time decay factor corresponding to the user's historical behavior, each historical interest vector can be weighted and averaged to obtain a comprehensive user interest vector.

[0087] S205. Based on a preset intent recognition classifier and the user interest vector, determine the intent recognition result of the current query information; the intent recognition result includes: intent category and confidence level.

[0088] It's important to note that the intent recognition classifier can comprehensively analyze the user's interest vector and the current query information through an attention mechanism to obtain the intent category and confidence level of the current query. Compared to directly identifying the intent of the current query, the intent recognition classifier can adaptively identify the intent of the current query based on the characteristics of the user's interest vector through the attention mechanism, thus improving the accuracy of identifying the user's current intent. The intent category can be the task intent required for intelligent scheduling, fault diagnosis, equipment maintenance, data sharing, etc., based on the retrieved data assets. The confidence level represents the probability of the current query information under each intent category.

[0089] For example, the user interest vector and the current query information are input into a pre-established intent recognition classifier. The intent recognition classifier can identify the actual task intent (such as intelligent scheduling, fault diagnosis, equipment operation and maintenance, data sharing, etc.) that the data assets to be retrieved by the current query information need to perform based on the characteristics of the user interest vector. The classifier outputs the intent category and confidence level associated with the user interest vector as the intent recognition result of the current query information.

[0090] S206. Determine the target retrieval result from the associated retrieval results using the user interest vector and intent recognition result.

[0091] The technical solution of this embodiment can determine the user interest vector based on the user's historical behavior and a preset time decay weight. This comprehensively considers the influence of user behavior at different times on the determination of user interests, improving the accuracy of determining the user interest vector. Furthermore, based on a preset intent recognition classifier and the user interest vector, the intent category and confidence level of the current query information are determined. By combining an attention mechanism with user interests, the task intent of the current query information can be identified, improving the accuracy of intent recognition and thus enhancing the personalization and accuracy of data retrieval.

[0092] Based on the above embodiments, this embodiment also provides an optional embodiment, which can further optimize the above embodiments and may further include:

[0093] The target search results are sorted to determine the target search data list;

[0094] The intent recognition classifier is optimized based on user behavior data related to the target retrieval data list.

[0095] The target retrieval data list can be displayed through a pre-set data retrieval display interface. Behavioral data can include user actions such as clicking, long-staying, and exiting the target retrieval data list, which are used to optimize the intent recognition classifier.

[0096] For example, after determining the target retrieval results, a pre-defined ranking model can be used to estimate a score relevant to the current user for each retrieved data asset (such as an estimated click-through rate). Then, the data assets in the target retrieval results are ranked according to the scores output by the ranking model to determine the target retrieval data list, which is then displayed through a pre-set data retrieval display interface. Furthermore, it can receive all user behavior data on the target retrieval data list displayed on the interface, and determine strong positive samples, strong negative samples, and uncertain samples based on the user's behavior data. This allows for the optimization of the intent recognition classifier based on each sample. Strong positive samples can be the intent category corresponding to the data asset where the user performs a click behavior, strong negative samples can be the intent category corresponding to the data asset where the user performs a bounce behavior, and uncertain samples can be the intent category corresponding to the data asset where the user performs a long dwell behavior.

[0097] Example 3

[0098] Figure 3 This is a schematic diagram of the structure of a data retrieval device provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes:

[0099] The acquisition module 31 can be used to acquire the current query information input by the user;

[0100] Matching module 32 can be used to semantically match the current query information with a pre-established standard database to determine the initial search results;

[0101] The association determination module 33 can be used to determine the association search results based on the initial search results and the preset semantic relationship graph;

[0102] The intent recognition module 34 can be used to determine the user interest vector and intent recognition result based on the user's historical behavior and the current query information;

[0103] The filtering model 35 can be used to determine the target retrieval result from the associated retrieval results through the user interest vector and intent recognition results.

[0104] The technical solution implemented in this device can obtain the user's current query information, perform semantic matching between the current query information and a pre-established standard database to determine the initial search results, and then determine related search results based on the initial search results and a preset semantic relationship graph. This solves the problem of isolated user query results and expands related search results. Then, based on the user's historical behavior and the current query information, user interest vectors and intent recognition results are determined. Using these user interest vectors and intent recognition results, target search results are determined from the related search results. This solves the problem that traditional retrieval systems cannot proactively understand the user's true needs, enabling the identification of the user's query intent and improving the accuracy of data retrieval and user experience.

[0105] Optionally, the method for establishing the standard database may include:

[0106] Acquiring multi-source heterogeneous data asset sets

[0107] Multimodal feature encoding is performed on the data asset set to obtain the corresponding semantic information dataset;

[0108] The standard database is constructed based on the semantic information dataset.

[0109] Optionally, the matching module 32 may include: a semantic encoding unit, a semantic constraint unit, and a semantic matching unit.

[0110] The semantic encoding unit can be used to semantically encode the current query information using a preset semantic encoding model to obtain the semantic vector of the current query information.

[0111] The semantic constraint unit can be used to perform semantic constraints on the current query information based on knowledge graph technology, so as to obtain the restricted semantic subspace of the current query information.

[0112] The semantic matching unit can be used to perform semantic matching in the pre-established standard database based on the restricted semantic subspace and semantic vector of the current query information to determine the initial retrieval result.

[0113] Optionally, the semantic matching unit can be used to recall a candidate set in the standard database based on the restricted semantic subspace of the current query information;

[0114] The semantic vector of the current query information is compared with the candidate set to calculate the similarity, and the initial retrieval result is determined based on the calculation result.

[0115] Optionally, the semantic relationship graph includes at least two nodes and at least one edge; the nodes are used to represent data assets; the edges are used to represent the associations between different nodes.

[0116] Correspondingly, the association determination module 33 can be specifically used to determine the target node in the semantic relationship graph based on the initial search results;

[0117] Based on the target node and the edges connected to the target node, determine the associated nodes;

[0118] The data assets corresponding to the target node and the associated node are identified as the associated retrieval results.

[0119] Optionally, the intent recognition module 34 can be used to determine the user interest vector based on the user's historical behavior and a preset time decay weight.

[0120] Based on a preset intent recognition classifier and the user interest vector, the intent recognition result of the current query information is determined; the intent recognition result includes: intent category and confidence level.

[0121] Optionally, the device further includes: an optimization module, which can be used to sort the target retrieval results to determine a target retrieval data list; and optimize the intent recognition classifier based on user behavior data related to the target retrieval data list.

[0122] The data retrieval device provided in the embodiments of the present invention can execute the data retrieval method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0123] Example 4

[0124] Figure 4A schematic diagram of an electronic device 40 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0125] like Figure 4 As shown, the electronic device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42 or a random access memory (RAM) 43, communicatively connected to the at least one processor 41. The memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes based on the computer program stored in the ROM 42 or loaded from storage unit 48 into the RAM 43. The RAM 43 may also store various programs and data required for the operation of the electronic device 40. The processor 41, ROM 42, and RAM 43 are interconnected via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.

[0126] Multiple components in electronic device 40 are connected to I / O interface 45, including: input unit 46, such as keyboard, mouse, etc.; output unit 47, such as various types of monitors, speakers, etc.; storage unit 48, such as disk, optical disk, etc.; and communication unit 49, such as network card, modem, wireless transceiver, etc. Communication unit 49 allows electronic device 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0127] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as data retrieval methods.

[0128] In some embodiments, the data retrieval method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the data retrieval method described above may be performed. Alternatively, in other embodiments, processor 41 may be configured to perform the data retrieval method by any other suitable means (e.g., by means of firmware).

[0129] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0130] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0131] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0132] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0133] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0134] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0135] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0136] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A data retrieval method, characterized by, The method comprises the following steps: obtaining current query information input by a user; performing semantic matching on the current query information and a pre-established standard database to determine an initial retrieval result; determining an associated retrieval result based on the initial retrieval result and a preset semantic relationship graph; determining a user interest vector and an intent recognition result according to a user historical behavior and the current query information; determining a target retrieval result from the associated retrieval result through the user interest vector and the intent recognition result.

2. The method of claim 1, wherein, The method for establishing the standard database comprises the following steps: obtaining a multi-source heterogeneous data asset set; performing multi-modal feature coding on the data asset set to obtain a corresponding semantic information data set; constructing the standard database based on the semantic information data set.

3. The method of claim 1, wherein, The method for performing semantic matching on the current query information and the pre-established standard database to determine an initial retrieval result comprises the following steps: performing semantic coding on the current query information through a preset semantic coding model to obtain a semantic vector of the current query information; performing semantic constraint on the current query information based on knowledge graph technology to obtain a restricted semantic subspace of the current query information; performing semantic matching on the pre-established standard database based on the restricted semantic subspace and the semantic vector of the current query information to determine the initial retrieval result.

4. The method of claim 3, wherein, The method for performing semantic matching on the pre-established standard database based on the restricted semantic subspace and the semantic vector of the current query information to determine the initial retrieval result comprises the following steps: recalling a candidate set in the standard database according to the restricted semantic subspace of the current query information; performing similarity calculation on the semantic vector of the current query information and the candidate set, and determining the initial retrieval result according to the calculation result.

5. The method of claim 1, wherein, The semantic relationship graph comprises at least two nodes and at least one edge; the nodes are used for representing data assets; and the edges are used for representing the association between different nodes. The method for determining an associated retrieval result based on the initial retrieval result and a preset semantic relationship graph comprises the following steps: determining a target node in the semantic relationship graph based on the initial retrieval result; determining an associated node according to the target node and an edge connected to the target node; determining data assets corresponding to the target node and the associated node as the associated retrieval result.

6. The method of claim 1, wherein, The method for determining a user interest vector and an intent recognition result according to a user historical behavior and the current query information comprises the following steps: determining a user interest vector according to a user historical behavior and a preset time decay weight; determining an intent recognition result of the current query information based on a preset intent recognition classifier and the user interest vector; the intent recognition result comprises an intent category and a confidence degree.

7. The method of claim 6, wherein, The method further comprises the following steps: sorting the target retrieval result to determine a target retrieval data list; optimizing the intent recognition classifier based on behavior data of a user with respect to the target retrieval data list.

8. A data retrieval apparatus, characterized by comprising: The method comprises the following steps: an obtaining module, configured to obtain current query information input by a user; a matching module, configured to perform semantic matching on the current query information and a pre-established standard database to determine an initial retrieval result; The association determining module is configured to determine the associated search result based on the initial search result and a preset semantic relation graph; The intention identifying module is configured to determine a user interest vector and an intention identifying result according to the user historical behavior and the current query information; The screening model is configured to determine a target search result from the associated search result based on the user interest vector and the intention identifying result.

9. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected with the at least one processor in communication; wherein The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the data search method in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the data search method in any one of claims 1-7 when executed.