Recall Method, Device, Equipment and Medium for Structured Data

By combining keyword recall and semantic recall methods, multiple recall strategies and semantic filtering technology are adopted to solve the problem of low accuracy of structured data recall, and more efficient and accurate data retrieval is achieved.

CN117093601BActive Publication Date: 2025-07-25BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311117801.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-31
Publication Date
2025-07-25
Estimated Expiration
2043-08-31

AI Technical Summary

Technical Problem

In the prior art, the recall accuracy of structured data is low and it is difficult to meet the query needs of users.

Method used

By combining keyword recall and semantic recall methods, multiple recall strategies are used to obtain structured data, and combined with semantic filtering technology, the accuracy and relevance of recall are improved.

Benefits of technology

Improve the efficiency and accuracy of structured data retrieval to ensure that the query results are more in line with user expectations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117093601B_ABST
    Figure CN117093601B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, device and medium for recalling structured data, relating to the field of artificial intelligence, and particularly to the field of data processing. By determining a structured data table corresponding to a query text in a structured database, where the structured data table contains multiple pieces of structured data; obtaining multiple types of keywords corresponding to the query text, and performing multi-way recall on the structured data in the structured data table based on a multi-way recall strategy corresponding to each type of keyword to obtain a first structured data set; performing semantic recall on the query text and the structured data in the structured data table, and performing semantic filtering on the structured data obtained by the semantic recall to obtain a second structured data set generated after filtering; obtaining a target structured data set according to the first structured data set and the second structured data set. This application combines keyword recall and semantic recall, which can improve the efficiency and accuracy of structured data retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of artificial intelligence and data processing, and in particular to the field of artificial intelligence. Specifically, it relates to a method, apparatus, device and medium for recalling structured data. Background Art

[0002] Structured data refers to data stored in a clearly defined and standardized format, which is applicable to a variety of different scenarios. It usually exists in the form of a table and has predefined fields and data types. For example, in the public security scenario, structured data includes data such as people, trains, airplanes, hotels, cases, and police situations, and these data are stored in a database in the form of a relational table. In related technologies, most use keyword search technology to recall relevant structured data from the database. However, since the matching information may be scattered in different field names and field values, the accuracy of this solution is relatively low. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, device and medium for recalling structured data.

[0004] According to one aspect of the present disclosure, there is provided a method for recalling structured data. By determining a structured data table corresponding to a query text in a structured database, where the structured data table contains multiple pieces of structured data; obtaining multiple types of keywords corresponding to the query text, and performing multi-way recall on the structured data in the structured data table based on a multi-way recall strategy corresponding to each type of keyword to obtain a first set of structured data; performing semantic recall on the query text and the structured data in the structured data table, and performing semantic filtering on the structured data obtained by the semantic recall to obtain a second set of structured data generated after filtering; and obtaining a target set of structured data according to the first set of structured data and the second set of structured data.

[0005] In the method for recalling structured data provided in this application, the keyword recall uses a multi-way recall strategy, which can more accurately find structured data related to the query, improve the accuracy and relevance of the query results. The semantic recall can consider a wider context and semantic relevance, enhance the relevance of the retrieval results, and combined with semantic filtering, can further screen and filter the recalled structured data to ensure that the results are more in line with user expectations. This application combines keyword recall and semantic recall, which can improve the efficiency and accuracy of structured data retrieval.

[0006] According to another aspect of the present disclosure, there is provided a recall device for structured data, including a determination module configured to determine a structured data table corresponding to a query text in a structured database, where the structured data table contains multiple pieces of structured data; a keyword recall module configured to obtain multiple types of keywords corresponding to the query text, and perform multi-way recall on the structured data in the structured data table based on a multi-way recall strategy corresponding to each type of keyword to obtain a first structured data set; a semantic recall module configured to perform semantic recall on the query text and the structured data in the structured data table, and perform semantic filtering on the structured data obtained by the semantic recall to obtain a second structured data set generated after filtering; and a data acquisition module configured to obtain a target structured data set according to the first structured data set and the second structured data set.

[0007] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned recall method for structured data.

[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the above-mentioned recall method for structured data.

[0009] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements the above-mentioned recall method for structured data when executed by a processor.

[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0012] Figure 1 is a schematic diagram of an exemplary embodiment of a recall method for structured data shown in the present application.

[0013] Figure 2 is a schematic diagram of an exemplary embodiment of another recall method for structured data shown in the present application.

[0014] Figure 3It is a schematic diagram of an exemplary embodiment of another structured data recall method shown in this application.

[0015] Figure 4 It is a schematic diagram of the architecture of a structured data recall method shown in this application.

[0016] Figure 5 It is a schematic diagram of a structured data recall device shown in this application.

[0017] Figure 6 It is a schematic diagram of an electronic device according to an exemplary embodiment of the present disclosure. Detailed implementation manners

[0018] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following.

[0019] Artificial Intelligence (AI) is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.). It has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include several aspects such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0020] Data processing: Data is a form of expression of facts, concepts, or instructions and can be processed by manual or automated devices. After data is interpreted and given a certain meaning, it becomes information. Data processing is the collection, storage, retrieval, processing, transformation, and transmission of data. The basic purpose of data processing is to extract and derive valuable and meaningful data for certain specific people from a large amount of data that may be chaotic and difficult to understand.

[0021] Figure 1 It is a schematic diagram of an exemplary embodiment of a structured data recall method shown in this application. As Figure 1 shown, the structured data recall method includes the following steps:

[0022] S101, determine the structured data table corresponding to the query text in the structured database, where the structured data table contains multiple pieces of structured data.

[0023] Receive the text input by the user as the query text, and according to the content and semantics of the query text, use relevant algorithms or rules to match the association between the query text and the structured database table.

[0024] For example, it can be matched by comparing the keywords in the query text with the field names, table names, table descriptions, etc. of the database table, or natural language processing techniques such as entity recognition and relation extraction can be used to understand the intention of the query text to determine the corresponding structured data table in the structured database.

[0025] S102, obtain multiple types of keywords corresponding to the query text, and perform multi-way recall on the structured data in the structured data table based on the multi-way recall strategy corresponding to each type of keyword to obtain the first structured data set.

[0026] Extract multiple types of keywords from the query text. Common keyword extraction methods include word segmentation, part-of-speech tagging, and named entity recognition, etc. For each keyword type, different recall strategies can be formulated to obtain relevant structured data. According to the recall strategies of different keyword types, execute these strategies in sequence and obtain relevant structured data. The obtained structured data is used as the first structured data, and all the first structured data together form the first structured data set.

[0027] S103, perform semantic recall on the query text and the structured data in the structured data table, and perform semantic filtering on the structured data obtained by the semantic recall to obtain the second structured data set generated after filtering.

[0028] Input the above query text and multiple pieces of structured data included in the structured data table into the trained semantic recall model, obtain multiple pieces of structured data output by the semantic recall model, and input the multiple pieces of structured data output by the semantic recall model and the above query text into the trained semantic filtering model together to obtain multiple pieces of structured data output by the semantic filtering model. The multiple pieces of structured data output by the semantic filtering model are used as the second structured data, and all the second structured data together form the second structured data set.

[0029] Optionally, the semantic recall model uses a large language pre-trained model as the base and combines a two-tower deep neural network (DNN). The semantic recall model can utilize the semantic representation ability of the pre-trained model and model and evaluate the similarity between the query text and multiple pieces of structured data through the two-tower DNN to achieve text recall. Among them, the large language pre-trained model can be the pre-trained language model (Bidirectional Encoder Representations from Transformers, BERT); among them, the two-tower DNN adopts a pairwise training method, where one tower processes the query text and the other tower processes multiple pieces of structured data contained in the structured data table. Each tower contains multiple layers of deep neural networks for extracting the feature representation of the input text. The outputs of the two towers are combined through a connection layer, and the similarity score between them is calculated.

[0030] Optionally, the semantic filtering model uses a large language pre-trained model as the base and combines a single-tower deep neural network (DNN). The semantic filtering model can utilize the semantic representation ability of the pre-trained model and match the scores of the query text and multiple pieces of structured data output by the semantic recall model through the single-tower DNN to achieve filtering. Among them, the large language pre-trained model can be the pre-trained language model (Bidirectional Encoder Representations from Transformers, BERT); among them, the single-tower DNN adopts a pointwise training method, that is, each query text and the corresponding structured data output by the semantic recall model are processed as independent training samples. For each training sample, the model classifies or regresses through a deep neural network to output the predicted matching score. Among them, when training the semantic filtering model, some counterexamples of multiple pieces of structured data recalled by the semantic recall model can be constructed as partial negative samples.

[0031] S104, obtain a target structured data set according to the first structured data set and the second structured data set.

[0032] Merge the first structured data set and the second structured data set to obtain a merged structured data set generated after merging; sort all the structured data contained in the merged structured data set, and obtain the top P structured data arranged after sorting as the target structured data; obtain a target structured data set composed of all the target structured data.

[0033] Among them, when sorting all the structured data included in the merged structured data set, a heuristic strategy or a CTR prediction model based on click logs can be used.

[0034] An embodiment of the present application proposes a method for retrieving structured data. By determining the structured data table corresponding to the query text in the structured database, where the structured data table contains multiple pieces of structured data; obtaining multiple types of keywords corresponding to the query text, and performing multi-way retrieval on the structured data in the structured data table based on the multi-way retrieval strategy corresponding to each type of keyword to obtain a first structured data set; performing semantic retrieval on the query text and the structured data in the structured data table, and performing semantic filtering on the structured data obtained by the semantic retrieval to obtain a second structured data set generated after filtering; and obtaining a target structured data set according to the first structured data set and the second structured data set. In the present application, the keyword retrieval uses a multi-way retrieval strategy, which can more accurately find the structured data related to the query, improve the accuracy and relevance of the query results. The semantic retrieval can consider a wider context and semantic relevance, enhance the relevance of the retrieval results, and combined with semantic filtering, can further screen and filter the retrieved structured data to ensure that the results are more in line with the user's expectations. The present application combines keyword retrieval with semantic retrieval, which can improve the efficiency and accuracy of structured data retrieval.

[0035] Figure 2 is a schematic diagram of an exemplary embodiment of another method for retrieving structured data shown in the present application, as Figure 2 shown, the method for retrieving structured data includes the following steps:

[0036] S201, determine the structured data table corresponding to the query text in the structured database, where the structured data table contains multiple pieces of structured data.

[0037] Obtain the query text input by the user. Optionally, the query text can be obtained through a user interface, an Application Programming Interface (API) request, or any other applicable means.

[0038] Perform intent recognition on the query text. Among them, the goal of intent recognition is to determine the main intent or purpose of the user's query. When performing intent recognition on the query text, a trained classification model can be used. The model receives the query text as input and outputs one or more intent categories. According to the intent recognition result, determine the structured data table corresponding to the query text in the structured database, which lays a foundation for subsequent structured data retrieval.

[0039] S202. Extract the attributes of the query text and obtain the extracted keywords as attribute keywords.

[0040] Optionally, when extracting the attributes of the query text, preprocessing can be performed on the query text, including removing stop words, punctuation marks, and other irrelevant characters, as well as performing operations such as word segmentation.

[0041] Attributes are usually specific information related to the query text, such as location, time, person, product, etc. In this application, a pre-trained model or rules can be used for entity recognition and relationship extraction, or a sequence labeling method can be adopted for attribute extraction to obtain the extracted keywords as attribute keywords.

[0042] S203. Obtain the remaining query text in the query text except for the attribute keywords, and perform importance analysis on the remaining query text to determine important keywords and unimportant keywords.

[0043] Further, use text processing techniques (such as regular expressions or natural language processing libraries) to remove the attribute keywords from the query text, and leave the remaining part as the remaining query text.

[0044] Perform importance analysis on the remaining query text to determine which keywords are important and related to the query intent, and which keywords are less important. Based on the results of the importance analysis, divide the keywords in the query text into important and unimportant keywords.

[0045] S204. Use the attribute keywords, important keywords, and unimportant keywords as multiple types of keywords.

[0046] S205. And perform multi-way recall on the structured data in the structured data table based on the multi-way recall strategy corresponding to each type of keyword to obtain the first structured data set.

[0047] Determine the multi-way recall strategy corresponding to each type of keyword, where the multi-way recall strategy includes at least two matching strategies among exact match, fuzzy match, regular match, and numerical comparison.

[0048] For each type of keyword, perform multi-way recall on the structured data in the structured data table according to the multi-way recall strategy to obtain the first structured data recalled by each path corresponding to this type of keyword. Generate the first structured data set according to all the first structured data corresponding to all types of keywords. In this application, based on the multi-way recall strategy, each recall strategy analyzes and recalls the data from different perspectives. This can increase the coverage of the recall and improve the recall rate, and can capture more potential matching items, thus providing more comprehensive results.

[0049] S206. Perform semantic recall on the query text and the structured data in the structured data table, and perform semantic filtering on the structured data obtained by semantic recall to obtain a second set of structured data generated after filtering.

[0050] Input the above query text and multiple pieces of structured data included in the structured data table into the trained semantic recall model, obtain multiple pieces of structured data output by the semantic recall model, and jointly input the multiple pieces of structured data output by the semantic recall model and the above query text into the trained semantic filtering model to obtain multiple pieces of structured data output by the semantic filtering model. Take the multiple pieces of structured data output by the semantic filtering model as the second structured data, and all the second structured data together form the second set of structured data.

[0051] Optionally, the semantic recall model adopts a large language pre-trained model base + pairwise two-tower deep neural networks (DNN).

[0052] Optionally, the semantic filtering model adopts a large language pre-trained model base + pointwise single-tower deep neural networks (DNN). When training the semantic filtering model, some counterexamples of multiple pieces of structured data recalled by the semantic recall model can be constructed as partial negative samples.

[0053] S207. Obtain a target set of structured data based on the first set of structured data and the second set of structured data.

[0054] As an implementable method, obtain the structured data that exists in both the first set of structured data and the second set of structured data as the target structured data; obtain a target set of structured data composed of all the target structured data. Since these data appear in both data sets, compare the two data sets and extract the data that exists in both as the target structured data. This method can ensure that the extracted data has higher accuracy and reliability, and can increase the integrity of the data, avoiding the loss of important information that may exist in one of the data sets.

[0055] As another implementable approach, the first structured data set and the second structured data set are merged to obtain a merged structured data set generated after the merge; all the structured data included in the merged structured data set are sorted, and the top P structured data after sorting are obtained as the target structured data; a target structured data set composed of all the target structured data is obtained. Among them, when sorting all the structured data included in the merged structured data set, a heuristic strategy or a ctr prediction model based on click logs can be used. This approach merges the first structured data set and the second structured data set, which can obtain a larger data set, thereby increasing the data volume and coverage, providing more comprehensive information. Sorting all the structured data in the merged structured data set can, according to specific sorting rules or algorithms, rank the most relevant or most valuable data at the front to improve the quality and relevance of the target structured data.

[0056] In the embodiments of the present application, keyword recall and semantic recall are combined. Keyword recall uses a multi-channel recall strategy, which can more accurately find structured data related to the query, improve the accuracy and relevance of the query results. Semantic recall can consider a wider context and semantic relevance, enhance the relevance of the retrieval results, and combined with semantic filtering, can further screen and filter the recalled structured data to ensure that the results are more in line with user expectations and improve the efficiency and accuracy of structured data retrieval.

[0057] Figure 3 is a schematic diagram of an exemplary implementation manner of another structured data recall method shown in the present application, as Figure 3 shown, this structured data recall method includes the following steps:

[0058] S301, determine the structured data table corresponding to the query text in the structured database, where the structured data table contains multiple pieces of structured data.

[0059] S302, obtain multiple types of keywords corresponding to the query text, and perform multi-channel recall on the structured data in the structured data table based on the multi-channel recall strategy corresponding to each type of keyword to obtain a first structured data set.

[0060] Regarding the specific implementation manners of steps S301 to S302, reference may be made to the specific introduction of the relevant parts in the above embodiments, and details will not be elaborated here.

[0061] S303, obtain a first query text vector corresponding to the query text, and obtain a candidate text vector corresponding to each piece of structured data in the structured data table.

[0062] Obtain the text vector corresponding to the query text as the first query text vector.

[0063] Obtain the text vector corresponding to each piece of structured data in the structured data table as the candidate text vector. Among them, each piece of structured data corresponds to one candidate text vector.

[0064] S304, obtain the vector similarity between the first query text vector and each candidate text vector.

[0065] S305, sort the vector similarities from largest to smallest, and according to the sorting result, determine the top N vector similarities as the target similarities.

[0066] S306, use the structured data corresponding to the target similarity as the initial second structured data.

[0067] S307, obtain the text vector to be filtered corresponding to each piece of initial second structured data, and obtain the second query text vector corresponding to the query text.

[0068] Obtain the text vector corresponding to each piece of initial second structured data as the text vector to be filtered.

[0069] Obtain the text vector corresponding to the query text as the second query text vector.

[0070] Among them, both the first query text vector and the second query text vector are vectors corresponding to the query text, but they can be obtained through different natural language processing techniques and models, so the first query text vector and the second query text vector are not exactly the same.

[0071] S308, obtain the matching score between the second query text vector and each text vector to be filtered.

[0072] Match the second query text vector with each text vector to be filtered, and obtain the matching score corresponding to each of the second query text vector and each text vector to be filtered.

[0073] S309, sort the matching scores from largest to smallest, and according to the sorting result, determine the top M matching scores as the target matching scores.

[0074] It is not difficult to understand that the value of M is less than the value of N.

[0075] S310, use the initial second structured data corresponding to the target matching score as the second structured data obtained after filtering.

[0076] S311, obtain the target structured data set according to the first structured data set and the second structured data set.

[0077] For the specific implementation of step S311, reference may be made to the specific introduction of the relevant part in the above embodiments, and details will not be elaborated herein.

[0078] In the embodiments of the present application, keyword recall is combined with semantic recall. The keyword recall uses a multi-channel recall strategy, which can more accurately find structured data related to the query, improve the accuracy and relevance of the query results. The semantic recall can consider a wider context and semantic relevance, enhance the relevance of the retrieval results, and combined with semantic filtering, can further screen and filter the recalled structured data to ensure that the results are more in line with user expectations and improve the efficiency and accuracy of structured data retrieval.

[0079] Figure 4 is a schematic architecture diagram of a method for recalling structured data shown in the present application. As Figure 4 shown, intent recognition is performed on the query text. According to the intent recognition result, the structured data table corresponding to the query text in the structured database is determined, laying a foundation for subsequent structured data recall.

[0080] In Term recall, attribute extraction is performed on the query text to obtain the extracted keywords as attribute keywords. The remaining query text in the query text is obtained, and importance parsing is performed on the remaining query text to determine important keywords and unimportant keywords. The attribute keywords, important keywords, and unimportant keywords are used as multiple types of keywords. And multi-channel recall is performed on the structured data in the structured data table based on the multi-channel recall strategy corresponding to each type of keyword to obtain a first structured data set.

[0081] In semantic recall, the above query text and multiple pieces of structured data included in the structured data table are input into the trained semantic recall model to obtain multiple pieces of structured data output by the semantic recall model. The multiple pieces of structured data output by the semantic recall model and the above query text are jointly input into the trained semantic filtering model to obtain multiple pieces of structured data output by the semantic filtering model. The multiple pieces of structured data output by the semantic filtering model are used as the second structured data, and all the second structured data together form a second structured data set.

[0082] The first structured data set and the second structured data set are merged to obtain a merged structured data set generated after merging; all the structured data included in the merged structured data set are sorted to obtain the top P structured data arranged after sorting as the target structured data; a target structured data set composed of all the target structured data is obtained.

[0083] Figure 5It is a schematic diagram of a recall device for structured data shown in this application. As Figure 5 shown, the recall device 500 for structured data includes a determination module 501, a keyword recall module 502, a semantic recall module 503, and a data acquisition module 504, where:

[0084] The determination module 501 is configured to determine the structured data table corresponding to the query text in the structured database, where the structured data table contains multiple pieces of structured data.

[0085] The keyword recall module 502 is configured to obtain multiple types of keywords corresponding to the query text, and perform multi-way recall on the structured data in the structured data table based on the multi-way recall strategy corresponding to each type of keyword, so as to obtain a first structured data set.

[0086] The semantic recall module 503 is configured to perform semantic recall on the query text and the structured data in the structured data table, and perform semantic filtering on the structured data obtained by the semantic recall, so as to obtain a second structured data set generated after filtering.

[0087] The data acquisition module 504 is configured to obtain a target structured data set according to the first structured data set and the second structured data set.

[0088] In this device, the keyword recall uses a multi-way recall strategy, which can more accurately find the structured data related to the query, improve the accuracy and relevance of the query results. The semantic recall can consider a wider context and semantic relevance, improve the relevance of the retrieval results, and combined with semantic filtering, can further screen and filter the recalled structured data to ensure that the results are more in line with the user's expectations. This application combines keyword recall and semantic recall, which can improve the efficiency and accuracy of structured data retrieval.

[0089] Further, the keyword recall module 502 is further configured to: extract attributes from the query text, and obtain the extracted keywords as attribute keywords; obtain the remaining query text in the query text except the attribute keywords, and perform importance analysis on the remaining query text to determine important keywords and unimportant keywords; use the attribute keywords, important keywords, and unimportant keywords as multiple types of keywords.

[0090] Further, the keyword recall module 502 is further configured to: determine a multi-channel recall strategy corresponding to each type of keyword, where the multi-channel recall strategy includes at least two of exact matching, fuzzy matching, regular matching, and numerical comparison; for each type of keyword, perform multi-channel recall on the structured data in the structured data table according to the multi-channel recall strategy, and obtain the first structured data recalled by each channel corresponding to the type of keyword; generate a first structured data set according to all the first structured data corresponding to all types of keywords.

[0091] Further, the semantic recall module 503 is further configured to: obtain a first query text vector corresponding to the query text, and obtain a candidate text vector corresponding to each piece of structured data in the structured data table; obtain the vector similarity between the first query text vector and each candidate text vector; sort the vector similarities from largest to smallest, and according to the sorting result, determine the top N vector similarities as the target similarities; use the structured data corresponding to the target similarities as the initial second structured data, and perform semantic filtering on the initial second structured data to obtain a second structured data set generated after filtering.

[0092] Further, the semantic recall module 503 is further configured to: obtain a text vector to be filtered corresponding to each piece of initial second structured data, and obtain a second query text vector corresponding to the query text; obtain the matching score between the second query text vector and each text vector to be filtered; sort the matching scores from largest to smallest, and according to the sorting result, determine the top M matching scores as the target matching scores; use the initial second structured data corresponding to the target matching scores as the second structured data obtained after filtering.

[0093] Further, the data acquisition module 504 is further configured to: obtain the structured data that exists in both the first structured data set and the second structured data set as the target structured data; obtain a target structured data set composed of all the target structured data.

[0094] Further, the data acquisition module 504 is further configured to: merge the first structured data set and the second structured data set to obtain a merged structured data set generated after merging; sort all the structured data included in the merged structured data set, and obtain the top P structured data after sorting as the target structured data; obtain a target structured data set composed of all the target structured data.

[0095] Further, the determination module 501 is further configured to: obtain the query text input by the user; perform intent recognition on the query text, and according to the intent recognition result, determine the structured data table corresponding to the query text in the structured database.

[0096] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0097] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0098] Figure 6 FIG. shows a schematic block diagram of an exemplary electronic device 600 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0099] As Figure 6 shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0100] A plurality of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0101] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the method for retrieving structured data. For example, in some embodiments, the method for retrieving structured data can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method for retrieving structured data described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the method for retrieving structured data in any other suitable manner (e.g., by means of firmware).

[0102] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0103] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0104] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0105] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0106] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0107] A computer system can include a client and a server. The client and the server are generally far apart from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.

[0108] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0109] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A recall method for structured data, comprising: Determine the structured data table corresponding to the query text in the structured database, wherein the structured data table contains multiple pieces of structured data; Obtain multiple types of keywords corresponding to the query text, and perform multi-way recall on the structured data in the structured data table based on the multi-way recall strategies corresponding to each type of keyword to obtain a first structured data set, wherein the multiple types of keywords include attribute keywords, important keywords, and unimportant keywords, and the multi-way recall strategies include at least two matching strategies among exact matching, fuzzy matching, regular matching, and numerical comparison; Perform semantic recall on the query text and the structured data in the structured data table, and perform semantic filtering on the structured data obtained by semantic recall to obtain a second structured data set generated after filtering; Obtain a target structured data set according to the first structured data set and the second structured data set.

2. The method according to claim 1, wherein The obtaining multiple types of keywords corresponding to the query text includes: Perform attribute extraction on the query text to obtain the extracted keywords as attribute keywords; Obtain the remaining query text except the attribute keywords in the query text, and perform importance analysis on the remaining query text to determine important keywords and unimportant keywords.

3. The method according to claim 2, wherein, The performing multi-way recall on the structured data in the structured data table based on the multi-way recall strategies corresponding to each type of keyword to obtain a first structured data set includes: Determine the multi-way recall strategies corresponding to each type of keyword; For each type of keyword, perform multi-way recall on the structured data in the structured data table according to the multi-way recall strategy to obtain the first structured data recalled by each path corresponding to this type of keyword; Generate the first structured data set according to all the first structured data corresponding to all types of keywords.

4. The method according to claim 1 or 3, wherein, The performing semantic recall on the query text and the structured data in the structured data table, and performing semantic filtering on the structured data obtained by semantic recall to obtain a second structured data set generated after filtering includes: Obtain a first query text vector corresponding to the query text, and obtain a candidate text vector corresponding to each piece of structured data in the structured data table; Obtain the vector similarity between the first query text vector and each candidate text vector; Sort the vector similarities from large to small, and according to the sorting result, determine the top N vector similarities as target similarities; Use the structured data corresponding to the target similarities as the initial second structured data, and perform semantic filtering on the initial second structured data to obtain a second structured data set generated after filtering.

5. The method according to claim 4, wherein The performing semantic filtering on the initial second structured data to obtain a second structured data set generated after filtering includes: Obtain a text vector to be filtered corresponding to each piece of the initial second structured data, and obtain a second query text vector corresponding to the query text; Obtain the matching scores corresponding to the second query text vector and each of the to-be-filtered text vectors; Sort the matching scores from largest to smallest, and based on the sorting result, determine the top M matching scores as the target matching scores; Use the initial second structured data corresponding to the target matching scores as the second structured data obtained after filtering.

6. The method according to claim 1 or 5, wherein, The obtaining of the target structured data set according to the first structured data set and the second structured data set includes: Obtain the structured data that exists in both the first structured data set and the second structured data set as the target structured data; Obtain the target structured data set composed of all the target structured data.

7. The method according to claim 1 or 5, wherein The obtaining of the target structured data set according to the first structured data set and the second structured data set includes: Merge the first structured data set and the second structured data set to obtain the merged structured data set generated after the merge; Sort all the structured data included in the merged structured data set, and obtain the top P structured data after sorting as the target structured data; Obtain the target structured data set composed of all the target structured data.

8. The method according to claim 1, wherein The determining of the structured data table corresponding to the query text in the structured database includes: Obtain the query text input by the user; Perform intent recognition on the query text, and based on the intent recognition result, determine the structured data table corresponding to the query text in the structured database.

9. A recall device for structured data, comprising: A determination module, configured to determine the structured data table corresponding to the query text in the structured database, where the structured data table contains multiple pieces of structured data; A keyword recall module, configured to obtain multiple types of keywords corresponding to the query text, and perform multi-way recall on the structured data in the structured data table based on the multi-way recall strategies corresponding to each type of keyword, so as to obtain a first structured data set, where the multiple types of keywords include attribute keywords, important keywords, and non-important keywords, and the multi-way recall strategies include at least two matching strategies of exact matching, fuzzy matching, regular matching, and numerical comparison; A semantic recall module, configured to perform semantic recall on the query text and the structured data in the structured data table, and perform semantic filtering on the structured data obtained by the semantic recall, so as to obtain a second structured data set generated after filtering; A data acquisition module, configured to obtain a target structured data set according to the first structured data set and the second structured data set.

10. The device according to claim 9, wherein, The keyword recall module is further configured to: Extract attributes from the query text, and obtain the extracted keywords as attribute keywords; Obtain the remaining query text in the query text except the attribute keywords, and perform importance analysis on the remaining query text to determine important keywords and non-important keywords.

11. The device according to claim 10, wherein, The keyword recall module is further configured to: Determine the multi-way recall strategies corresponding to each type of keyword; For each type of keyword, perform multi-channel recall on the structured data in the structured data table according to the multi-channel recall strategy, and obtain the first structured data recalled by each channel corresponding to this type of keyword; Generate the first structured data set according to all the first structured data corresponding to all types of keywords.

12. The device according to claim 9 or 11, wherein The semantic recall module is further configured to: Obtain the first query text vector corresponding to the query text, and obtain the candidate text vector corresponding to each piece of structured data in the structured data table; Obtain the vector similarity corresponding to the first query text vector and each candidate text vector; Sort the vector similarities from largest to smallest, and according to the sorting result, determine the top N vector similarities as the target similarities; Use the structured data corresponding to the target similarities as the initial second structured data, and perform semantic filtering on the initial second structured data to obtain the second structured data set generated after filtering.

13. The device according to claim 12, wherein, The semantic recall module is further configured to: Obtain the text vector to be filtered corresponding to each piece of the initial second structured data, and obtain the second query text vector corresponding to the query text; Obtain the matching score corresponding to the second query text vector and each text vector to be filtered; Sort the matching scores from largest to smallest, and according to the sorting result, determine the top M matching scores as the target matching scores; Use the initial second structured data corresponding to the target matching scores as the second structured data obtained after filtering.

14. The device according to claim 9 or 13, wherein The data acquisition module is further configured to: Obtain the structured data that exists in both the first structured data set and the second structured data set as the target structured data; Obtain the target structured data set composed of all the target structured data.

15. The device according to claim 9 or 13, wherein The data acquisition module is further configured to: Merge the first structured data set and the second structured data set to obtain the merged structured data set generated after merging; Sort all the structured data included in the merged structured data set, and obtain the top P structured data after sorting as the target structured data; Obtain the target structured data set composed of all the target structured data.

16. The device according to claim 9, wherein, The determination module is further configured to: Obtain the query text input by the user; Perform intent recognition on the query text, and according to the intent recognition result, determine the structured data table corresponding to the query text in the structured database.

17. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

19. A computer program product comprising a computer program which, when executed by a processor, implements the steps of the method according to any one of claims 1 - 8.

Citation Information

Patent Citations

  • Searching method and device based on keywords and semantics, equipment and storage medium

    CN115438166A

  • Searching method and device, model training method and device, electronic equipment and storage medium

    CN116662633A