Search methods, devices, storage media and electronic equipment
By using intent recognition and entity mapping tables in the human-computer interaction system, the accuracy and recall issues under diverse command formats are solved, achieving efficient and accurate information retrieval.
Patent Information
- Application Number
- CN202311465848.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2043-11-06
AI Technical Summary
Existing human-computer interaction systems face challenges in terms of accuracy and recall, especially when faced with diverse user command formats. Traditional methods require a large amount of data for training, consume significant computing resources, and have slow data update and iteration.
By identifying the professional field of user commands through intent recognition, a corresponding entity word mapping table is established for precise retrieval, and entity words are expanded when necessary. Combined with word vectors and contextual features for composite intent recognition, the accuracy and recall of retrieval are improved.
It improves the efficiency and accuracy of retrieval, reduces the demand for computing resources, and ensures the freshness and rapid updating of data.
Smart Images

Figure CN117633169B_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification pertain to the field of human-computer interaction, and particularly relate to retrieval methods, devices, storage media, and electronic devices. Background Technology
[0002] Human-Computer Interaction (HCI) refers to the process and technology of information exchange and interaction between human users and computer systems. It aims to interpret user questions as queries and retrieve the most relevant answers from structured or unstructured data.
[0003] In natural language-based human-computer interaction systems, accurate understanding of user questions, efficient and accurate retrieval of answers, and timeliness of data are crucial issues. However, significant challenges remain in improving the accuracy and recall of human-computer interaction systems, necessitating solutions with higher accuracy and better timeliness. Summary of the Invention
[0004] This specification provides a retrieval method, apparatus, storage medium, and electronic device, the technical solutions of which are as follows:
[0005] Firstly, embodiments of this specification provide a retrieval method, including:
[0006] Obtain user instructions;
[0007] The user instruction is subjected to intent recognition to obtain the intent recognition result of the user instruction, and an entity word mapping table corresponding to the intent recognition result is obtained;
[0008] Entity word recognition is performed on the user command to obtain the entity word set of the user command;
[0009] The entity word set is queried based on the entity word mapping table, and the query result is used as the retrieval result.
[0010] Secondly, embodiments of this specification provide a retrieval device, including:
[0011] The instruction acquisition unit is configured to acquire user instructions;
[0012] The mapping table acquisition unit is configured to perform intent recognition on the user instruction, obtain the intent recognition result of the user instruction, and acquire the entity word mapping table corresponding to the intent recognition result;
[0013] An entity word recognition unit is configured to perform entity word recognition on the user instruction to obtain a set of entity words for the user instruction;
[0014] The query unit is configured to query the entity word set based on the entity word mapping table and use the query result as the retrieval result.
[0015] Thirdly, embodiments of this specification provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect above.
[0016] Fourthly, embodiments of this specification provide an electronic device, including:
[0017] One or more processors, and memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method described in the first aspect above.
[0018] The beneficial effects of the technical solutions provided by one or more embodiments of this specification include at least the following:
[0019] First, user commands are identified through intent recognition to determine their professional domain, thus narrowing the search to that specific domain. This allows for the acquisition of an entity mapping table relevant to that domain, which boasts high collectability of entity words, ample training data, and effective training, thereby improving the efficiency and accuracy of subsequent searches. Second, for the limited-sample interaction environment of user commands, a composite intent recognition method based on word vectors and contextual features is introduced during intent recognition, enhancing its accuracy. Third, during the search process, a precise query is performed first. If the precise query fails to yield valid results, the entity words are expanded for a secondary query, further improving the recall rate. Fourth, the entity mapping table is stored and updated, improving search efficiency while maintaining data freshness. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this specification, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a diagram illustrating an exemplary human-computer interaction application scenario.
[0022] Figure 2 This is a schematic flowchart illustrating the retrieval method provided in the embodiments of this specification.
[0023] Figure 3This is a flowchart illustrating the intent identification process in the retrieval method provided in the embodiments of this specification.
[0024] Figure 4 This is a schematic block diagram of a retrieval device provided in an embodiment of this specification.
[0025] Figure 5 A schematic block diagram of another retrieval device provided in the embodiments of this specification.
[0026] Figure 6 This is a schematic block diagram of an electronic device provided as an embodiment of this specification. Detailed Implementation
[0027] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings.
[0028] The terms "first," "second," "third," etc., in the description, claims, and accompanying drawings herein are used to distinguish different objects and not to describe a particular order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus. Depending on the context, the word "if," as it applies herein, can be interpreted as "when," "when," "in response to determination," or "in response to detection."
[0029] like Figure 1 The diagram illustrates an application scenario for a Human-Computer Interaction (HCI) system. HCI refers to the process and technology of information exchange and interaction between human users and computer systems. It aims to understand user questions as queries and retrieve the most relevant answers from structured or unstructured data. With advancements in natural language processing and artificial intelligence, the performance and effectiveness of HCI systems have continuously improved, leading to their widespread application in various fields and scenarios, such as intelligent assistants, search engines, education, and healthcare.
[0030] Text retrieval is a crucial component of human-computer interaction systems. Most current text retrieval systems employ natural language-based similarity matching algorithms, which primarily retrieve answers by calculating the similarity between texts. However, these algorithms suffer from problems such as low retrieval efficiency, poor recall, and untimely data preservation.
[0031] In view of this, the embodiments of this specification provide a new retrieval method and apparatus that, while having better retrieval efficiency and accuracy, also maintains the freshness of the data.
[0032] Figure 2 This is a flowchart illustrating the retrieval method provided in the embodiments of this specification. This method can be implemented by a corresponding retrieval device. The method may include the following steps:
[0033] Step 202: Obtain user instructions;
[0034] Step 204: Perform intent recognition on the user command to obtain the intent recognition result of the user command, and obtain the entity word mapping table corresponding to the intent recognition result;
[0035] Step 206: Perform entity word recognition on the user command to obtain the entity word set of the user command;
[0036] Step 208: Query the entity word set based on the entity word mapping table, and use the query result as the retrieval result.
[0037] The above user instructions are user-input instructions. Depending on the way the user interacts with the system and the platform used, the instructions can also exist in various forms: (1) Keyword query: Users can use keywords or phrases to describe the information they need. For example, users can enter "weather forecast" to get the weather information for the day; (2) Question query: Users ask questions in the form of questions. For example, users can ask "Will it rain tomorrow?" to inquire about the weather tomorrow; (3) Command instructions: Users use specific commands to tell the system to perform a certain operation. For example, users can enter "open album" to start the album application; (4) Natural language instructions: Users use natural language to communicate. It can be a complete sentence or dialogue. For example, users can say "I want to find a nearby Italian restaurant" to find nearby restaurants; (5) Voice instructions: Users use voice input devices, such as voice assistants or voice recognition technology, to interact with the system through voice. For example, users can say "play a Beatles song" to request that music be played.
[0038] These forms can be used individually or in combination to meet user needs and interaction methods. With the development of human-computer interaction technology, the forms of user instructions are constantly being enriched and innovated; therefore, the embodiments in this specification do not limit the specific form of user instructions.
[0039] However, in human-computer dialogue, the diversity of command formats presents greater challenges to text retrieval. The current common practice is to train models on large datasets of user commands to learn as many command expressions as possible and improve retrieval accuracy. However, on the one hand, training large-scale models typically requires massive amounts of data to achieve sufficient generalization ability; collecting, cleaning, and labeling large-scale data is a time-consuming and complex process. Furthermore, storing and processing large-scale data also requires significant computing and storage resources. On the other hand, as the amount of data continues to increase, the model's training and iteration times will further lengthen, thus impacting the model's deployment schedule.
[0040] This specification proposes, in its embodiments, to perform intent recognition on user commands before text retrieval. Then, based on the intent recognition results, the professional field to which the user command belongs is determined. This is equivalent to an initial screening of user commands, narrowing the scope of text retrieval to different professional fields and reducing the computational load. Furthermore, this specification proposes establishing entity word mapping tables for different professional fields to provide data support for text retrieval. After determining the professional field to which the user command belongs, precise retrieval is performed based on the corresponding field entity word mapping table, resulting in better accuracy and recall. Simultaneously, entity words from different professional fields are highly collectible, with ample training corpora and good training effects. Compared to training with extensive big data models, this approach offers greater operability in query matching and facilitates updates and iterations, resulting in higher data freshness.
[0041] Traditional intent recognition, in the field of natural language processing, refers to the process of determining a user's intent or purpose by analyzing their linguistic input. However, the intent recognition in the embodiments of this specification differs from traditional intent recognition in several ways. Firstly, in human-computer interaction systems, users often express their intent with a limited amount of language, resulting in insufficient context. Secondly, users often express their intent casually, making it difficult to determine whether they expressed their exact intent at the beginning, middle, or end of the conversation. Furthermore, the mutual influence between different intents further contributes to ambiguity. Additionally, user commands may contain entity information, such as proper nouns, which conventional methods cannot recognize, thus treating them like other text and causing biases in intent recognition.
[0042] In view of this, embodiments of this specification provide the following method for intent recognition. For example... Figure 3 As shown, the method includes the following steps:
[0043] Step 302: Convert the user instruction to obtain a set of word vectors for the user instruction;
[0044] Step 304: Convert the user instruction to obtain the context feature set of the user instruction;
[0045] Step 306: Convert the word vector set and the context feature set to obtain the composite feature set of the user instruction;
[0046] Step 308: Normalize the composite features in the composite feature set, and take the composite feature with the largest value after normalization as the intention recognition result of the user instruction.
[0047] Before performing intent recognition, the dialogue text must first be converted into a form that a computer can process. This process requires feature extraction from the dialogue text. Currently, various feature extraction models exist, such as Word2Vec or BERT. However, the features extracted in this way often only contain shallow, single semantic information. In the context of human-computer dialogue, where text volume is limited, the extracted features are clearly insufficient to clearly express the user's intent. Therefore, to better reflect the user's intent in human-computer interaction systems, this specification proposes an intent recognition method that integrates contextual features in its embodiments.
[0048] In step 302, firstly, the user command Query is segmented to obtain the segmentation set Q = {Q1, Q2, Q3, ..., Q...} n}; Entity word extraction is performed on the word segmentation set Q to obtain the entity word representation Q of Q. E Then, Q is encoded. E Encode the commands to obtain a set of word vectors for the user's instructions.
[0049] In specific embodiments described in this specification, the encoder may be based on an attention mechanism for Q. E Encoding is performed. Attention mechanisms primarily simulate the human ability to allocate attention when processing information. They can help assign different weights and importance to different parts of the input when processing related data, thereby better focusing on key information. Based on attention mechanisms, Q... E Encoding can dynamically generate word vectors based on different contexts in user instructions.
[0050] Building upon the steps described above, and considering that the arbitrariness of user commands in form can lead to unclear intent, step 304 further considers obtaining intent from the temporal features of the dialogue. A recurrent neural network is used to capture these temporal features to obtain the specific intent within the dialogue. Based on the aforementioned characteristics of recurrent neural networks, in a specific embodiment of this specification, the user command is first converted into a text sequence. This text sequence is then input into a recurrent neural network to extract contextual features, resulting in a set of contextual features for the user command.
[0051] In step 306, after obtaining the word vector set and context feature set of the user instruction, in order to further highlight the information in the user instruction, it is necessary to fuse the word vector set and the context feature set to obtain a composite feature set. This composite feature set contains both the entity information of the original instruction and the features of the context, making the features that reflect the user's true intent more prominent.
[0052] In a specific embodiment of this specification, the weights of different context features in the context feature set can be determined first, and then the context features are inserted into the word vector set based on these weights. For example, if the user's instruction is "What are some good Sichuan restaurants nearby?", then "nearby" and "Sichuan cuisine" will have higher weights than other entity words. During the fusion of the word vector set and the context feature set, "nearby" and "Sichuan cuisine" will be made more prominent based on their weights. In a specific embodiment of this specification, inserting context features into the word vector set can be done through concatenation.
[0053] Finally, in step 308, after obtaining the composite feature set, the intent recognition result of the user instruction can be calculated from it. Specifically, the composite features in the composite feature set are normalized, and the composite feature with the largest normalized value is taken as the intent recognition result of the user instruction. In a specific embodiment of this specification, the softmax function is preferred for calculating the composite feature normalization. The softmax function converts a set of real numbers into a vector representing a probability distribution, such that each element is between 0 and 1, and the sum of all elements is 1. The softmax function exponentializes the input and then normalizes the exponentialized result, so that the value of each element represents the probability of the corresponding category. The largest value corresponds to the highest category probability, and the smaller the value, the lower the probability. Using this method, the result of user instruction intent recognition is more accurate, laying a better foundation for subsequent information retrieval.
[0054] After normalizing the composite feature set, the composite feature with the largest value is taken as the intent recognition result of the user command, thus classifying the user command into a specific professional field. For example, if the user command is "What are some good Sichuan restaurants nearby?", after normalization, "nearby + eat" is found to be the user's intent, and the user command can be identified as belonging to the catering field.
[0055] In the embodiments of this specification, after determining the professional field to which the user instruction belongs, it is necessary to further obtain the entity word mapping table corresponding to the professional field. The purpose of the entity word mapping table is to store and manage entity words. Optionally, the entity words can be converted into a discrete representation, which makes them easier to process in a computer.
[0056] In the embodiments of this specification, the construction of the entity word mapping table includes the following steps:
[0057] Entity word collection: First, entity words need to be collected from corpora, text data, or other sources. Entity words can be specific entity identifiers such as person names, place names, organization names, and product names.
[0058] Assign a field: Assign a unique field to each entity term. The field can be an integer, a string, or other format.
[0059] Constructing the mapping table: Create a mapping table for each entity term and its corresponding field. This can be achieved using dictionaries, hash tables, or other data structures.
[0060] In this embodiment, entity words are first collected from internet data sources through data cleaning. Then, fields are assigned to these entity words according to different domains, resulting in entity word mapping tables for different professional domains. For example, an entity word mapping table for the catering industry can be established, with fields such as city name, street location, restaurant name, average price per person, and cuisine. Different entity word mapping tables can be established based on different professional domains. This embodiment does not further limit the specific form of the entity word mapping table. Finally, a mapping table is constructed for each entity word and its corresponding fields, which can be implemented using a dictionary, hash table, or other data structures.
[0061] In specific embodiments of this specification, the entity term mapping table is stored as an entity. The storage format includes, but is not limited to, NAS, distributed file system, or database. By storing the entity as an entity, the efficiency of retrieval can be improved, and the entity term mapping table can be easily updated to ensure the freshness of the data.
[0062] Based on the results of user intent recognition, the corresponding entity word mapping table can be obtained, which can reduce the workload of subsequent retrieval, improve retrieval efficiency, make the retrieval more targeted, and ensure higher accuracy and recall.
[0063] After obtaining the entity word mapping table, it is necessary to further obtain the entity words contained in the user commands. The goal of entity word recognition is to identify entities with specific meanings from text, such as names of people, places, organizations, dates, and times. The key to entity word recognition is annotation, which involves labeling each word in the input text sequence as either an entity or a non-entity. Common entity categories include names of people, places, organizations, dates, and times.
[0064] In the embodiments of this specification, a trained natural language model is used to identify entity words in user commands. The main goal of the natural language model is to model the probability distribution of language, that is, to predict the probability of the next word based on the preceding words or context. This modeling can be used for various tasks, such as language generation, text classification, machine translation, and speech recognition. With the development of deep learning, the performance and application scope of natural language models are constantly improving, bringing higher accuracy and flexibility to natural language processing tasks. In the embodiments of this specification, no specific limitations are made on the training process of the natural language model.
[0065] Finally, the identified entity words are queried in the entity word mapping table. In this embodiment, a search technique is used to search the entity word mapping table, replacing the traditional method of matching and searching using an algorithm model. In a specific embodiment of this specification, the Zsearch search platform can be used to query entity words.
[0066] When no matching result is found in the entity term mapping table, a blank value is returned as the search result. To improve the recall rate, in a specific embodiment of this specification, the entity term set is expanded to obtain a mapping term set. This expansion mainly uses synonyms, near-synonyms, and aliases. For example, "Chengdu" has aliases such as Rongcheng, Jincheng, Furongcheng, Jinguancheng, and Tianfuzhiguo. Therefore, when no matching result is found in the entity term mapping table, "Chengdu" can be expanded to obtain a mapping term set of "Rongcheng, Jincheng, Furongcheng, Jinguancheng, and Tianfuzhiguo". This mapping term set is then queried again in the entity term mapping table, and the query result is used as the search result.
[0067] By using the methods described above, the scope of user commands can be broadened to include all entity content that the user commands may contain, thereby further improving the retrieval rate.
[0068] Since aliases, synonyms, and near-synonyms of entity words are updated over time, the entity word mapping table can be updated in specific embodiments of this specification. Updates can be passive, occurring at regular intervals, or proactive, triggered by specific events. For network data, which iterates rapidly, updates can be performed daily. When significant events occur, such as changes in company management, the entity word mapping table can be proactively updated. This specification does not further limit the specific update cycle or events described in the embodiments.
[0069] In one embodiment of this specification, a retrieval device is also included. This retrieval device can be deployed on a local computer, a cloud server, or a mobile device. Its main function is to receive user instructions and convert these instructions into search results to present to the user. The user instructions may be requests received from a client, which may be made using the HTTP protocol or other communication protocols.
[0070] Please see Figure 4 The embodiment shown in this specification provides a retrieval device. For example... Figure 4 As shown, the retrieval device 400 includes an instruction acquisition unit 402, a mapping table acquisition unit 404, an entity word recognition unit 406, and a query unit 408. The main functions of each component are as follows:
[0071] The instruction acquisition unit 402 is configured to acquire user instructions.
[0072] The mapping table acquisition unit 404 is configured to perform intent recognition on the user instruction, obtain the intent recognition result of the user instruction, and acquire an entity word mapping table corresponding to the intent recognition result.
[0073] The entity word recognition unit 406 is configured to perform entity word recognition on the user instruction to obtain the entity word set of the user instruction.
[0074] The query unit 408 is configured to query the entity word set based on the entity word mapping table and use the query result as the retrieval result.
[0075] The user instructions described above are user-inputted instructions. Depending on the interaction method between the user and the system and the platform used, these instructions can take various forms, such as keyword searches and question queries. These forms can be used individually or in combination to meet user needs and interaction methods. With the development of human-computer interaction technology, the forms of user instructions are constantly being enriched and innovated; therefore, the embodiments in this specification do not limit the specific form of the user instructions.
[0076] However, in human-computer dialogue, the diversity of command formats presents greater challenges to text retrieval. The current common practice is to train models on large datasets of user commands to learn as many command expressions as possible and improve retrieval accuracy. However, on the one hand, training large-scale models typically requires massive amounts of data to achieve sufficient generalization ability; collecting, cleaning, and labeling large-scale data is a time-consuming and complex process. Furthermore, storing and processing large-scale data also requires significant computing and storage resources. On the other hand, as the amount of data continues to increase, the model's training and iteration times will further lengthen, thus impacting the model's deployment schedule.
[0077] This specification proposes, in its embodiments, to perform intent recognition on user commands before text retrieval, and then conduct targeted searches based on the results of intent recognition. This is equivalent to a preliminary screening of user commands, narrowing the scope of text retrieval to different professional fields and effectively improving retrieval efficiency. Furthermore, this specification proposes to establish entity word mapping tables for different professional fields to provide a data foundation for text retrieval. By identifying user commands within a specific professional field and using the corresponding entity word mapping table for precise retrieval, accuracy and recall are better guaranteed. Entity words from different professional fields are highly collectible, with ample training corpora and good training results. Compared to training with extensive big data models, this approach offers greater operability in query matching and facilitates updates and iterations, resulting in higher data freshness.
[0078] Figure 5 This specification illustrates a mapping table acquisition unit 404 provided in an embodiment, which includes a vector calculation unit 502, a feature calculation unit 504, a feature compositing unit 506, an intent recognition unit 508, and a table retrieval unit 510. The main functions of each component are as follows:
[0079] The vector calculation unit 502 is configured to convert the user instruction to obtain a set of word vectors for the user instruction.
[0080] The feature calculation unit 504 is configured to convert the user instruction to obtain a set of context features of the user instruction.
[0081] The feature composite unit 506 is configured to transform the word vector set and the context feature set to obtain the composite feature set of the user instruction.
[0082] The intent recognition unit 508 is configured to perform normalization calculation on the composite features in the composite feature set, and take the composite feature with the largest value after normalization calculation as the intent recognition result of the user instruction.
[0083] The table retrieval unit 510 is configured to obtain the corresponding entity word mapping table based on the intent recognition result.
[0084] First, the user command Query is segmented into words to obtain the word segmentation set Q = {Q1, Q2, Q3, ..., Qn}. Entity words are extracted from the word segmentation set Q to obtain the entity word representation QE of Q. Then, QE is encoded by an encoder to obtain the word vector set of the user command.
[0085] In specific embodiments of this specification, the encoder can encode the QE based on an attention mechanism. An attention mechanism is a technique commonly used in machine learning and natural language processing to simulate the human ability to allocate attention when processing information. It helps the model assign different weights and importance to different parts of the input when processing sequential data, thereby better focusing on key information. Attention mechanisms are commonly used for sequence-to-sequence tasks, such as machine translation, text summarization, and speech recognition. The vector computation unit encodes the QE based on the attention mechanism and can dynamically generate word vectors according to different contexts in user instructions.
[0086] In a specific embodiment of this specification, the method further includes a weight determination subunit and a feature insertion subunit. The weight determination subunit is configured to determine the weight of each context feature in the context feature set. The feature insertion subunit is configured to insert the context features in the context feature set into the word vector set based on the weights, thereby obtaining a composite feature set of the user instruction. For example, if the user instruction is "What are some good Sichuan restaurants nearby?", then "nearby" and "Sichuan cuisine" will have higher weights than other entity words. In a specific embodiment of this specification, inserting context features into the word vector set can be done through concatenation.
[0087] Finally, after obtaining the composite feature set, the intent recognition result of the user instruction can be calculated from it. Specifically, the composite features in the composite feature set are normalized, and the composite feature with the largest normalized value is taken as the intent recognition result of the user instruction.
[0088] After determining the intent recognition result of the user command, it is equivalent to filtering the user command into a specific professional field, and then obtaining the entity word mapping table corresponding to the above professional field for precise retrieval.
[0089] In this embodiment of the specification, a mapping table generation unit is also included, configured to obtain the entity word mapping table by crawling and cleaning different categories of data from the network. For example, an entity word mapping table for company positions can be established, which can set fields such as company name, position, and person name; an entity word mapping table related to city catering can also be established, which can set city name, street name, and restaurant name, etc.; different entity word mapping tables can be established according to different professional fields, and this embodiment of the specification does not make specific limitations.
[0090] When no matching result is found in the entity term mapping table, a blank value is returned as the search result. To improve the recall rate, in a specific embodiment of this specification, the entity term set is expanded by an entity term expansion unit to obtain a mapping term set for the entity term set. The expansion mainly uses synonyms, near-synonyms, and aliases. For example, Chengdu is an alias of Rongcheng, Guangzhou is an alias of Yangcheng, chairman and CEO are near-synonyms, boss and boss are near-synonyms, etc. The query unit queries the entity term mapping set again in the entity term mapping table and uses the query result as the search result.
[0091] In specific embodiments of this specification, a mapping table storage unit is also included, configured to store the entity term mapping table via NAS, a distributed file system, or a database. The entity term mapping table is stored in a physical storage manner, which includes, but is not limited to, NAS, a distributed file system, or a database. Physical storage improves retrieval efficiency and facilitates updates to the entity term mapping table, ensuring data freshness.
[0092] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0093] Additionally, embodiments of this specification also provide another computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of any one of the methods described in the foregoing embodiments. If the constituent modules of the above-described apparatus are implemented as software functional units and sold or used as independent products, they can be stored in the computer-readable storage medium.
[0094] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this specification are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)).
[0095] This specification also provides an electronic device, including:
[0096] One or more processors, and
[0097] A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method described in any of the foregoing method embodiments.
[0098] This specification also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any one of the methods described in the foregoing method embodiments.
[0099] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. Unless otherwise specified, the technical features of this embodiment and its implementation can be combined arbitrarily.
[0100] in, Figure 6 An exemplary architecture of electronic device 600 is shown, which may specifically include: processor 610, disk drive 620, input / output interface 630, network interface 640, and memory 650. The processor 610, disk drive 620, input / output interface 630, network interface 640, and memory 650 can communicate with each other via a communication bus.
[0101] The processor 610 can be implemented using a general-purpose CPU, microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits to execute relevant programs and implement the technical solution provided in this application.
[0102] The memory 650 can be implemented in the form of ROM (Read Only Memory), RAM (Read Access Memory), static memory, dynamic storage devices, etc. The memory 650 can store the operating system 651 for controlling the operation of the electronic device 600, and the basic input / output system (BIOS) 652 for controlling the first-level operation of the electronic device 600. Additionally, it can store a web browser 653, a data storage management system 654, and a retrieval device 655, etc. In summary, when implementing the technical solution provided in this application through software or firmware, the relevant program code is stored in the memory 650 and is called and executed by the processor 610.
[0103] The input / output interface 630 is used to connect input / output modules to enable information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0104] The network interface 640 is used to connect the communication module (not shown in the figure) to enable communication and interaction between the device and other devices. The communication module can communicate via wired means (e.g., USB, Ethernet cable) or wireless means (e.g., mobile network, Wi-Fi, Bluetooth).
[0105] The bus includes a pathway for transmitting information between various components of the device (e.g., processor 610, disk drive 620, input / output interface 630, network interface 640, and memory 650).
[0106] It should be noted that although the above-described device only shows the processor 610, disk drive 620, input / output interface 630, network interface 640, memory 650, bus, etc., in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the method of this application, and does not necessarily include all the components shown in the figures.
[0107] The embodiments described above are merely preferred embodiments of this specification and are not intended to limit the scope of this specification. Any modifications and improvements made by those skilled in the art to the technical solutions of this specification without departing from the spirit of this specification should fall within the protection scope defined by the claims of this specification.
Claims
1. Retrieval methods, including: Obtain user instructions; The user instruction is subjected to intent recognition to obtain the intent recognition result of the user instruction, and an entity word mapping table corresponding to the intent recognition result is obtained; Entity word recognition is performed on the user command to obtain the entity word set of the user command; The entity word set is queried based on the entity word mapping table, and the query result is used as the retrieval result; The step of performing intent recognition on the user instruction to obtain the intent recognition result of the user instruction includes: The user instruction is converted to obtain a set of word vectors for the user instruction; The user instruction is converted to obtain a set of context features of the user instruction; The word vector set and the context feature set are transformed to obtain the composite feature set of the user instruction; The composite features in the composite feature set are normalized, and the composite feature with the largest normalized value is taken as the intent recognition result of the user instruction.
2. The method according to claim 1, wherein converting the user instruction to obtain the word vector set of the user instruction comprises: The user instructions are encoded and converted using an attention mechanism to obtain a set of word vectors for the user instructions.
3. The method according to claim 1, wherein converting the user instruction to obtain the context feature set of the user instruction includes: The user instructions are converted to obtain a text sequence of the user instructions; The text sequence is transformed using a recurrent neural network to obtain a set of contextual features for the user instruction.
4. The method according to claim 1, wherein converting the word vector set and the context feature set to obtain the composite feature set of the user instruction comprises: Determine the weight of each context feature in the context feature set; Based on the weights, the context features in the context feature set are inserted into the word vector set to obtain the composite feature set of the user instruction.
5. The method according to claim 1, comprising: The entity word mapping table is obtained by crawling and cleaning data of different categories from the network.
6. The method according to claim 1, comprising: The entity word mapping table is stored via NAS, a distributed file system, or a database, and is updated passively or actively.
7. The method according to claim 1, wherein performing entity word recognition on the user instruction to obtain the entity word set of the user instruction includes: The user command is input into a natural language model for conversion, resulting in a set of entity words for the user command.
8. The method according to claim 1, comprising: When the query result is a blank value, the entity word set is expanded to obtain the mapping word set of the entity word set.
9. The method of claim 8, comprising: The entity word mapping table is used to query the mapping word set, and the query result is used as the retrieval result.
10. A retrieval device, comprising: The instruction acquisition unit is configured to acquire user instructions; The mapping table acquisition unit is configured to perform intent recognition on the user instruction, obtain the intent recognition result of the user instruction, and acquire the entity word mapping table corresponding to the intent recognition result; An entity word recognition unit is configured to perform entity word recognition on the user instruction to obtain a set of entity words for the user instruction; The query unit is configured to query the entity word set based on the entity word mapping table, and use the query result as the retrieval result; The mapping table acquisition unit includes: A vector calculation unit is configured to convert the user instruction to obtain a set of word vectors for the user instruction; The feature calculation unit is configured to convert the user instruction to obtain a set of context features of the user instruction; The feature composite unit is configured to transform the word vector set and the context feature set to obtain the composite feature set of the user instruction; The intent recognition unit is configured to perform normalization calculation on the composite features in the composite feature set, and take the composite feature with the largest value after normalization calculation as the intent recognition result of the user instruction. The table retrieval unit is configured to obtain the corresponding entity word mapping table based on the intent recognition result.
11. The apparatus of claim 10, comprising: The vector calculation unit is configured to encode and convert the user instruction based on an attention mechanism to obtain a set of word vectors for the user instruction.
12. The apparatus of claim 10, comprising: A sequence generation unit is configured to convert the user instruction to obtain a text sequence of the user instruction; The feature calculation unit is configured to transform the text sequence based on a recurrent neural network to obtain a set of context features for the user instruction.
13. The apparatus of claim 10, wherein the feature composite unit comprises: The weight determination subunit is configured to determine the weight of each context feature in the context feature set; The feature insertion subunit is configured to insert context features from the context feature set into the word vector set based on the weights, thereby obtaining a composite feature set of the user instruction.
14. The apparatus of claim 10, comprising: The mapping table generation unit is configured to obtain the entity word mapping table by crawling and cleaning data of different categories from the network.
15. The apparatus of claim 10, comprising: Mapping table storage units are configured to be stored via NAS, distributed file systems, or databases, and updated in a passive or active manner.
16. The apparatus of claim 10, comprising: The entity word recognition unit is configured to convert the user command through a natural language model to obtain a set of entity words for the user command.
17. The apparatus of claim 10, comprising: The entity word expansion unit is configured to expand the entity word set when the query result is a blank value, so as to obtain a mapping word set of the entity word set.
18. The apparatus of claim 17, comprising: The query unit queries the mapping word set based on the entity word mapping table and uses the query result as the retrieval result.
19. A computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method according to any one of claims 1-9.
20. Electronic devices, including: One or more processors, and A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1-9.
Citation Information
Patent Citations
Question and answer processing method and device, computer equipment and storage medium
CN110427467A