Information query method and device based on section classification, electronic equipment and medium

By constructing a segmented and hierarchical information query method, and using a dual filtering method based on scenario-based vocabulary relationships and relevance scoring formulas, the problem of inaccurate information queries in the insurance and financial fields has been solved, achieving higher query accuracy.

CN116450916BActive Publication Date: 2026-01-06PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310430929.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-11
Publication Date
2026-01-06
Estimated Expiration
2043-04-11

AI Technical Summary

Technical Problem

Existing search engines often fail to provide accurate results for information searches in the insurance and finance sectors because the frequency of specialized terms is similar to that of common terms.

Method used

A segmented and hierarchical information query method is constructed. This method acquires scene vocabulary, performs relationship analysis, generates a reference thesaurus and expands the search engine, and uses a relevance scoring formula to calculate the relevance of the answer text. This double screening is then used to improve accuracy.

Benefits of technology

The dual screening mechanism significantly improves the accuracy of information retrieval, ensuring that the search results better meet user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116450916B_ABST
    Figure CN116450916B_ABST
Patent Text Reader

Abstract

The application relates to artificial intelligence and discloses an information query method based on fixed-section grading, which comprises the following steps: constructing a generated reference word library based on searched scene keywords and vocabulary relationships in multiple scene vocabularies, expanding an information search engine by using the reference word library to obtain a standard search engine; searching for answer texts corresponding to user query texts in an answer content library by using the standard search engine, and calculating the correlation score between the answer texts and the user query texts; performing primary screening on multiple answer text sets according to the correlation score to obtain an initial answer text set, and performing secondary screening on the initial answer text set according to a fixed-level rule constructed by using the reference word library to obtain a standard answer text set. In addition, the application also relates to blockchain technology, and the scene keywords can be stored in the nodes of the blockchain. The application further provides an information query device based on fixed-section grading, an electronic device and a storage medium. The application can improve the accuracy of information query.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and in particular to a method, apparatus, electronic device, and storage medium for information retrieval based on segmentation and hierarchical classification. Background Technology

[0002] With the widespread adoption of the internet, search engines have become increasingly prevalent in internet products and have become a universal tool for search scenarios. When using search engines to search for information, results are typically retrieved and ranked based on the relevance of text matching. While search engines offer flexible scoring methods for information retrieval, these scores are based on word frequency within the content database. In certain specific fields, such as insurance or finance, the frequency of specialized terms is close to that of general terms, leading to inaccurate results. Therefore, a more accurate information retrieval method is urgently needed. Summary of the Invention

[0003] This invention provides a method, apparatus, electronic device, and storage medium for information retrieval based on segmentation and hierarchical classification, with the main objective of improving the accuracy of information retrieval.

[0004] To achieve the above objectives, the present invention provides an information query method based on segmented and hierarchical classification, comprising:

[0005] Obtain multiple scenario terms in the target application scenario, search for scenario keywords in the multiple scenario terms, and perform relationship analysis on the multiple scenario terms to obtain the word relationships;

[0006] A reference lexicon is generated based on multiple scenario keywords and the word relationships. The reference lexicon is then used to expand the preset information search engine to obtain a standard search engine.

[0007] The standard search engine is used to search for one or more answer texts corresponding to the user's query text in a preset answer content library, and the relevance score between the answer text and the user's query text is calculated according to a preset relevance scoring formula;

[0008] The initial answer text set is obtained by initially screening multiple answer text sets based on the relevance score. The ranking rules are then constructed using the reference lexicon, and the initial answer text set is further screened based on the ranking rules to obtain the standard answer text set.

[0009] Optionally, the search for scene keywords among the multiple scene terms includes:

[0010] Extract multiple scene words from preset historical scene data and summarize the multiple scene words into a scene training set;

[0011] The keyword extraction model is obtained by training a pre-defined convolutional network model using the scene training set.

[0012] The keyword extraction model is used to extract keywords from the scene vocabulary to obtain scene keywords.

[0013] Optionally, the step of performing relational analysis on multiple scenario words to obtain word relationships includes:

[0014] Identify the word types corresponding to the scene words, and construct hierarchical or synonym relationships between scene words of the same word type;

[0015] Based on preset correspondence rules, construct correspondence relationships for scene words corresponding to different word types;

[0016] The hierarchical relationship, the synonym relationship, and the correspondence relationship are summarized into a lexical relationship.

[0017] Optionally, constructing the rating rules using the reference lexicon includes:

[0018] The vocabulary in the reference lexicon is parsed, and the scene keywords in the parsed vocabulary are taken as first-level vocabulary. The matching rule is set as the successful matching of the first-level vocabulary.

[0019] The words that conform to the initial rules are counted to obtain the number of words, and the initial ranking rules are constructed based on the interval in which the number of words falls.

[0020] The matching rules and the initial rating rules are combined to obtain the rating rules.

[0021] Optionally, the step of further filtering the initial answer text set according to the grading rules to obtain the standard answer text set includes:

[0022] Extract the initial answer texts with consistent relevance scores from the initial answer text set as the target text set;

[0023] The target texts in the target text set are classified according to the classification rules to obtain the standard answer text set.

[0024] Optionally, calculating the relevance score between the answer text and the user query text according to a preset relevance scoring formula includes:

[0025] The preset relevance scoring formula is as follows:

[0026]

[0027] Where score(D,Q) is the relevance score, Q is the user query text, D is the answer text, and q is the answer value. i For the i-th keyword in the answer text, IDF(q) i f(q) represents the inverse text frequency of keywords in the user query text, k1 and b are preset adjustment parameters, avgdl is the preset document average, |D| is the modulus corresponding to the answer text, and f(q) represents the inverse text frequency of keywords in the user query text. i ,) represents the i-th keyword q i Frequency of occurrence in answer text D.

[0028] Optionally, the step of extending the preset information search engine using the reference thesaurus to obtain a standard search engine includes:

[0029] The word to be added is obtained by using a preset information search engine to identify information from the reference word library;

[0030] The words to be added are added to the extended vocabulary of the word segmenter in the information search engine to obtain a standard search engine.

[0031] To address the above problems, the present invention also provides an information query device based on segmented and hierarchical classification, the device comprising:

[0032] The relation analysis module is used to acquire multiple scenario words in the target application scenario, search out scenario keywords in the multiple scenario words, and perform relation analysis on the multiple scenario words to obtain word relations;

[0033] The engine extension module is used to construct and generate a reference lexicon based on multiple scenario keywords and the word relationships, and to extend the preset information search engine using the reference lexicon to obtain a standard search engine;

[0034] The score calculation module is used to search for one or more answer texts corresponding to the user query text in the preset answer content library using the standard search engine, and calculate the relevance score between the answer text and the user query text according to the preset relevance scoring formula;

[0035] The dual-screening module is used to initially screen multiple answer text sets based on the relevance scores to obtain an initial answer text set, construct a grading rule using the reference lexicon, and further screen the initial answer text set based on the grading rule to obtain a standard answer text set.

[0036] To address the above problems, the present invention also provides an electronic device, the electronic device comprising:

[0037] At least one processor; and,

[0038] A memory communicatively connected to the at least one processor; wherein,

[0039] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the segmented and hierarchical information query method described above.

[0040] To address the aforementioned problems, the present invention also provides a storage medium storing at least one computer program, which is executed by a processor in an electronic device to implement the above-described segmented and hierarchical information query method.

[0041] In this embodiment of the invention, a reference thesaurus is generated by constructing multiple scenario keywords and vocabulary relationships. This reference thesaurus is then used to expand a preset information search engine, resulting in a standard search engine. Expanding the information search engine increases search accuracy. The standard search engine is then used to search a preset answer content database for one or more answer texts corresponding to the user's query text. These answer text sets are initially filtered based on relevance scores and then further filtered according to the grading rules to obtain a standard answer text set. The initial and subsequent filtering provide a dual filtering effect, improving the accuracy of information retrieval. Therefore, the information retrieval method, device, electronic device, and storage medium based on segmented grading proposed in this invention can solve the problem of low accuracy in information retrieval. Attached Figure Description

[0042] Figure 1 This is a flowchart illustrating an embodiment of the information query method based on segmentation and hierarchical classification provided by the present invention.

[0043] Figure 2 for Figure 1 A detailed implementation flowchart of one of the steps;

[0044] Figure 3 This is a functional block diagram of an information query device based on segmentation and hierarchical classification provided in an embodiment of the present invention;

[0045] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the segmented and hierarchical information query method according to an embodiment of the present invention.

[0046] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0047] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0048] This application provides a segmented and hierarchical information query method. The executing entity of this segmented and hierarchical information query method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the segmented and hierarchical information query method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0049] Reference Figure 1 The diagram shown is a flowchart illustrating an information query method based on segmented hierarchical classification according to an embodiment of the present invention. In this embodiment, the information query method based on segmented hierarchical classification includes the following steps S1-S4:

[0050] S1. Obtain multiple scenario terms under the target application scenario, search for scenario keywords among the multiple scenario terms, and perform relationship analysis on the multiple scenario terms to obtain the word relationships.

[0051] In this embodiment of the invention, the target application scenario can refer to business scenarios in different fields, such as business scenarios in the insurance field or the financial field. Scenario vocabulary refers to professional terms or operational terms of different types that appear in the specific operations of the business scenario.

[0052] For example, taking the insurance scenario as the target application scenario, there are two important categories of terms in the insurance context: insurance products and diseases. There are certain restrictions on which products can be insured for which diseases. There are also specialized terms in the insurance scenario, such as underwriting, claims, and reimbursement.

[0053] Specifically, refer to Figure 2 As shown, the scene keywords retrieved from the search for multiple scene terms include:

[0054] S11. Extract multiple scene words from preset historical scene data, and summarize the multiple scene words into a scene training set;

[0055] S12. The preset convolutional network model is trained using the scene training set to obtain the keyword extraction model;

[0056] S13. Extract keywords from the scene vocabulary according to the keyword extraction model to obtain scene keywords.

[0057] In detail, the preset convolutional network model includes a variety of different network structures.

[0058] Furthermore, the step of performing relational analysis on multiple scenario words to obtain word relationships includes:

[0059] Identify the word types corresponding to the scene words, and construct hierarchical or synonym relationships between scene words of the same word type;

[0060] Based on preset correspondence rules, construct correspondence relationships for scene words corresponding to different word types;

[0061] The hierarchical relationship, the synonym relationship, and the correspondence relationship are summarized into a lexical relationship.

[0062] In detail, the terminology type can be product type or symptom type. In the insurance scenario, different products have different versions or different levels. For example, the same insurance policy may be divided into different versions for adults, children, and adults. Different diseases may also have hierarchical or synonymous relationships. Therefore, hierarchical or synonymous relationships can be constructed between scenario terms corresponding to the same terminology type. Since there is also a correspondence between different diseases and different products, a correspondence relationship can be constructed for scenario terms corresponding to different terminology types according to preset correspondence rules.

[0063] S2. Based on multiple scenario keywords and the word relationships, a reference lexicon is generated. The reference lexicon is used to expand the preset information search engine to obtain a standard search engine.

[0064] In this embodiment of the invention, the step of constructing a reference lexicon based on multiple scenario keywords and the lexical relationships includes:

[0065] The pre-acquired thesaurus template is divided into regions to obtain keyword regions and relationship regions;

[0066] Multiple scenario keywords are stored in the keyword region, and the word relationships are stored in the relationship region to obtain the reference lexicon.

[0067] In detail, the reference lexicon is a comprehensive lexicon suitable for a specific application scenario.

[0068] Furthermore, the step of expanding the preset information search engine using the reference lexicon to obtain a standard search engine includes:

[0069] The word to be added is obtained by using a preset information search engine to identify information from the reference word library;

[0070] The words to be added are added to the extended vocabulary of the word segmenter in the information search engine to obtain a standard search engine.

[0071] In detail, the reference lexicon includes words of different categories, and words within the same category may have hierarchical, synonymous, or mutually exclusive relationships. A preset information search engine is used to identify information from the reference lexicon to obtain words to be added. These words are those that the information search engine can recognize and are added to the extended lexicon of the word segmenter in the information search engine, rather than adding all words from the maintained lexicon to the extended lexicon of the word segmenter in the information search engine.

[0072] The question-and-answer search engine mentioned above is ES (Elasticsearch, a distributed search engine), which has now become a common tool in search scenarios. It can rank search results based on scoring algorithms and the relevance of text matching.

[0073] S3. Using the standard search engine, search in the preset answer content library to obtain one or more answer texts corresponding to the user query text, and calculate the relevance score between the answer text and the user query text according to the preset relevance scoring formula.

[0074] In this embodiment of the invention, the standard search engine is used to search in a preset answer content library to obtain one or more answer texts corresponding to the user's query text. The preset answer content library contains multiple answers to common questions, so there are multiple different answers. The search yields one or more answer texts corresponding to the user's query text.

[0075] Specifically, the step of calculating the relevance score between the answer text and the user query text according to a preset relevance scoring formula includes:

[0076] The preset relevance scoring formula is as follows:

[0077]

[0078] Where score(D,Q) is the relevance score, Q is the user query text, D is the answer text, and q is the answer value. i For the i-th keyword in the answer text, IDF(q) if(q) represents the inverse text frequency of keywords in the user query text, k1 and b are preset adjustment parameters, avgdl is the preset document average, |D| is the modulus corresponding to the answer text, and f(q) represents the inverse text frequency of keywords in the user query text. i ,) represents the i-th keyword q i Frequency of occurrence in answer text D.

[0079] S4. Based on the relevance score, perform an initial screening of multiple answer text sets to obtain an initial answer text set. Use the reference lexicon to construct a grading rule, and perform a second screening of the initial answer text set based on the grading rule to obtain a standard answer text set.

[0080] In this embodiment of the invention, the initial screening of multiple answer text sets based on the relevance scores to obtain an initial answer text set includes:

[0081] The answer texts in the multiple answer text sets are sorted in descending order of relevance scores to obtain a text ranking list.

[0082] Select a predetermined number of answer texts that rank at the top of the text ranking list as the initial answer text set.

[0083] In detail, initial screening of multiple sets of answer texts can filter out a batch of answer texts that meet the requirements.

[0084] Specifically, the step of constructing classification rules using the reference lexicon includes:

[0085] The vocabulary in the reference lexicon is parsed, and the scene keywords in the parsed vocabulary are taken as first-level vocabulary. The matching rule is set as the successful matching of the first-level vocabulary.

[0086] The words that conform to the initial rules are counted to obtain the number of words, and the initial ranking rules are constructed based on the interval in which the number of words falls.

[0087] The matching rules and the initial rating rules are combined to obtain the rating rules.

[0088] In detail, let's take the insurance scenario as an example. The recall results are categorized based on the types of words in the user's query and the matching of specific words. For example, the most important product terms and disease terms are classified as primary terms; if a user's query contains a primary term, at least one must be matched. Different levels of recall results can also be assigned based on the number of primary terms matched in the user's query. Simultaneously, other professional terms in this field can be considered secondary terms. Combining the matching results of primary and secondary terms, the recall results can be further categorized.

[0089] Furthermore, the step of further filtering the initial answer text set according to the grading rules to obtain the standard answer text set includes:

[0090] Extract the initial answer texts with consistent relevance scores from the initial answer text set as the target text set;

[0091] The target texts in the target text set are classified according to the classification rules to obtain the standard answer text set.

[0092] In detail, the classification is based on the various types of words in the thesaurus and the relationships between them. When the matching results for these special words cannot be well classified, such as when there are many recall results in the same level and few special words in the user's query text, the text set at the same level can be segmented by combining the relevance score.

[0093] In this embodiment of the invention, a reference thesaurus is generated by constructing multiple scenario keywords and lexical relationships. This reference thesaurus is then used to expand a preset information search engine, resulting in a standard search engine. Expanding the information search engine increases search accuracy. The standard search engine is then used to search a preset answer content database for one or more answer texts corresponding to the user's query text. These answer text sets are initially filtered based on relevance scores and then further filtered according to the grading rules to obtain a standard answer text set. The initial and subsequent filtering provide a dual filtering effect, improving the accuracy of information retrieval. Therefore, the segmented and graded information retrieval method proposed in this invention can solve the problem of low accuracy in information retrieval.

[0094] like Figure 3 The diagram shown is a functional block diagram of an information query device based on segmentation and hierarchical classification provided in an embodiment of the present invention.

[0095] The segmented and hierarchical information query device 100 of this invention can be installed in an electronic device. Depending on the functions implemented, the segmented and hierarchical information query device 100 may include a relationship analysis module 101, an engine expansion module 102, a score calculation module 103, and a dual filtering module 104. The module described in this invention can also be called a unit, referring to a series of computer program segments that can be executed by the processor of an electronic device and perform a fixed function, stored in the memory of the electronic device.

[0096] In this embodiment, the functions of each module / unit are as follows:

[0097] The relationship analysis module 101 is used to acquire multiple scene words in the target application scenario, search out scene keywords in the multiple scene words, and perform relationship analysis on the multiple scene words to obtain word relationships.

[0098] The engine extension module 102 is used to construct and generate a reference lexicon based on multiple scenario keywords and the word relationships, and to extend the preset information search engine using the reference lexicon to obtain a standard search engine;

[0099] The score calculation module 103 is used to search for one or more answer texts corresponding to the user query text in the preset answer content library using the standard search engine, and calculate the relevance score between the answer text and the user query text according to the preset relevance scoring formula;

[0100] The dual screening module 104 is used to perform an initial screening of multiple answer text sets based on the relevance score to obtain an initial answer text set, construct a grading rule using the reference lexicon, and perform a second screening of the initial answer text set based on the grading rule to obtain a standard answer text set.

[0101] In detail, the specific implementation methods of each module of the segmented and hierarchical information query device 100 are as follows:

[0102] Step 1: Obtain multiple scenario terms in the target application scenario, search for scenario keywords in the multiple scenario terms, and perform relationship analysis on the multiple scenario terms to obtain the word relationships.

[0103] In this embodiment of the invention, the target application scenario can refer to business scenarios in different fields, such as business scenarios in the insurance field or the financial field. Scenario vocabulary refers to professional terms or operational terms of different types that appear in the specific operations of the business scenario.

[0104] For example, taking the insurance scenario as the target application scenario, there are two important categories of terms in the insurance context: insurance products and diseases. There are certain restrictions on which products can be insured for which diseases. There are also specialized terms in the insurance scenario, such as underwriting, claims, and reimbursement.

[0105] Specifically, the search for scene keywords among the multiple scene terms includes:

[0106] Extract multiple scene words from preset historical scene data and summarize the multiple scene words into a scene training set;

[0107] The keyword extraction model is obtained by training a pre-defined convolutional network model using the scene training set.

[0108] The keyword extraction model is used to extract keywords from the scene vocabulary to obtain scene keywords.

[0109] In detail, the preset convolutional network model includes a variety of different network structures.

[0110] Furthermore, the step of performing relational analysis on multiple scenario words to obtain word relationships includes:

[0111] Identify the word types corresponding to the scene words, and construct hierarchical or synonym relationships between scene words of the same word type;

[0112] Based on preset correspondence rules, construct correspondence relationships for scene words corresponding to different word types;

[0113] The hierarchical relationship, the synonym relationship, and the correspondence relationship are summarized into a lexical relationship.

[0114] In detail, the terminology type can be product type or symptom type. In the insurance scenario, different products have different versions or different levels. For example, the same insurance policy may be divided into different versions for adults, children, and adults. Different diseases may also have hierarchical or synonymous relationships. Therefore, hierarchical or synonymous relationships can be constructed between scenario terms corresponding to the same terminology type. Since there is also a correspondence between different diseases and different products, a correspondence relationship can be constructed for scenario terms corresponding to different terminology types according to preset correspondence rules.

[0115] Step 2: Based on multiple scenario keywords and the word relationships, construct a reference lexicon, and use the reference lexicon to expand the preset information search engine to obtain a standard search engine.

[0116] In this embodiment of the invention, the step of constructing a reference lexicon based on multiple scenario keywords and the lexical relationships includes:

[0117] The pre-acquired thesaurus template is divided into regions to obtain keyword regions and relationship regions;

[0118] Multiple scenario keywords are stored in the keyword region, and the word relationships are stored in the relationship region to obtain the reference lexicon.

[0119] In detail, the reference lexicon is a comprehensive lexicon suitable for a specific application scenario.

[0120] Furthermore, the step of expanding the preset information search engine using the reference lexicon to obtain a standard search engine includes:

[0121] The word to be added is obtained by using a preset information search engine to identify information from the reference word library;

[0122] The words to be added are added to the extended vocabulary of the word segmenter in the information search engine to obtain a standard search engine.

[0123] In detail, the reference lexicon includes words of different categories, and words within the same category may have hierarchical, synonymous, or mutually exclusive relationships. A preset information search engine is used to identify information from the reference lexicon to obtain words to be added. These words are those that the information search engine can recognize and are added to the extended lexicon of the word segmenter in the information search engine, rather than adding all words from the maintained lexicon to the extended lexicon of the word segmenter in the information search engine.

[0124] The question-and-answer search engine mentioned above is ES (Elasticsearch, a distributed search engine), which has now become a common tool in search scenarios. It can rank search results based on scoring algorithms and the relevance of text matching.

[0125] Step 3: Use the standard search engine to search for one or more answer texts corresponding to the user's query text in the preset answer content library, and calculate the relevance score between the answer text and the user's query text according to the preset relevance scoring formula.

[0126] In this embodiment of the invention, the standard search engine is used to search in a preset answer content library to obtain one or more answer texts corresponding to the user's query text. The preset answer content library contains multiple answers to common questions, so there are multiple different answers. The search yields one or more answer texts corresponding to the user's query text.

[0127] Specifically, the step of calculating the relevance score between the answer text and the user query text according to a preset relevance scoring formula includes:

[0128] The preset relevance scoring formula is as follows:

[0129]

[0130] Where score(D,Q) is the relevance score, Q is the user query text, D is the answer text, and q is the answer value. i For the i-th keyword in the answer text, IDF(q) i f(q) represents the inverse text frequency of keywords in the user query text, k1 and b are preset adjustment parameters, avgdl is the preset document average, |D| is the modulus corresponding to the answer text, and f(q) represents the inverse text frequency of keywords in the user query text. i ,) represents the i-th keyword q iFrequency of occurrence in answer text D.

[0131] Step 4: Perform an initial screening of multiple answer text sets based on the relevance scores to obtain an initial answer text set. Construct a grading rule using the reference lexicon, and further screen the initial answer text set according to the grading rule to obtain a standard answer text set.

[0132] In this embodiment of the invention, the initial screening of multiple answer text sets based on the relevance scores to obtain an initial answer text set includes:

[0133] The answer texts in the multiple answer text sets are sorted in descending order of relevance scores to obtain a text ranking list.

[0134] Select a predetermined number of answer texts that rank at the top of the text ranking list as the initial answer text set.

[0135] In detail, initial screening of multiple sets of answer texts can filter out a batch of answer texts that meet the requirements.

[0136] Specifically, the step of constructing classification rules using the reference lexicon includes:

[0137] The vocabulary in the reference lexicon is parsed, and the scene keywords in the parsed vocabulary are taken as first-level vocabulary. The matching rule is set as the successful matching of the first-level vocabulary.

[0138] The words that conform to the initial rules are counted to obtain the number of words, and the initial ranking rules are constructed based on the interval in which the number of words falls.

[0139] The matching rules and the initial rating rules are combined to obtain the rating rules.

[0140] In detail, let's take the insurance scenario as an example. The recall results are categorized based on the types of words in the user's query and the matching of specific words. For example, the most important product terms and disease terms are classified as primary terms; if a user's query contains a primary term, at least one must be matched. Different levels of recall results can also be assigned based on the number of primary terms matched in the user's query. Simultaneously, other professional terms in this field can be considered secondary terms. Combining the matching results of primary and secondary terms, the recall results can be further categorized.

[0141] Furthermore, the step of further filtering the initial answer text set according to the grading rules to obtain the standard answer text set includes:

[0142] Extract the initial answer texts with consistent relevance scores from the initial answer text set as the target text set;

[0143] The target texts in the target text set are classified according to the classification rules to obtain the standard answer text set.

[0144] In detail, the classification is based on the various types of words in the thesaurus and the relationships between them. When the matching results for these special words cannot be well classified, such as when there are many recall results in the same level and few special words in the user's query text, the text set at the same level can be segmented by combining the relevance score.

[0145] In this embodiment of the invention, a reference thesaurus is constructed by using multiple scenario keywords and word relationships. This reference thesaurus is then used to expand a preset information search engine, resulting in a standard search engine. Expanding the information search engine increases search accuracy. The standard search engine is then used to search a preset answer content database for one or more answer texts corresponding to the user's query text. These answer text sets are initially filtered based on relevance scores and then further filtered according to the grading rules to obtain a standard answer text set. The initial and subsequent filtering provide a dual filtering effect, improving the accuracy of information retrieval. Therefore, the segmented and graded information retrieval device proposed in this invention can solve the problem of low accuracy in information retrieval.

[0146] like Figure 4 The diagram shown is a structural schematic of an electronic device that implements a segmented and hierarchical information query method according to an embodiment of the present invention.

[0147] The electronic device 1 may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13. It may also include a computer program stored in the memory 11 and capable of running on the processor 10, such as an information query program based on segmentation and hierarchical classification.

[0148] In some embodiments, the processor 10 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control unit of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. It executes programs or modules stored in the memory 11 (e.g., executing segmented and hierarchical information query programs) and calls data stored in the memory 11 to perform various functions of the electronic device and process data.

[0149] The memory 11 includes at least one type of readable storage medium, including flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of an electronic device, such as a portable hard drive. In other embodiments, the memory 11 can be an external storage device of the electronic device, such as a plug-in portable hard drive, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. Furthermore, the memory 11 can include both internal and external storage units of the electronic device. The memory 11 can be used not only to store application software and various types of data installed on the electronic device, such as code for a segmented and hierarchical information query program, but also to temporarily store data that has been output or will be output.

[0150] The communication bus 12 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable communication between the memory 11 and at least one processor 10, etc.

[0151] The communication interface 13 is used for communication between the aforementioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, Bluetooth interface, etc.), typically used to establish communication connections between the electronic device and other electronic devices. The user interface may be a display, an input unit (such as a keyboard), or, optionally, a standard wired or wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the electronic device and to display a visual user interface.

[0152] Figure 4 Only electronic devices with components are shown; those skilled in the art will understand that... Figure 4 The structure shown does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0153] For example, although not shown, the electronic device may also include a power supply (such as a battery) to power the various components. Preferably, the power supply can be logically connected to the at least one processor 10 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.

[0154] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.

[0155] The information query program based on segmented hierarchy stored in the memory 11 of the electronic device 1 is a combination of multiple instructions, which, when run in the processor 10, can achieve the following:

[0156] Obtain multiple scenario terms in the target application scenario, search for scenario keywords in the multiple scenario terms, and perform relationship analysis on the multiple scenario terms to obtain the word relationships;

[0157] A reference lexicon is generated based on multiple scenario keywords and the word relationships. The reference lexicon is then used to expand the preset information search engine to obtain a standard search engine.

[0158] The standard search engine is used to search for one or more answer texts corresponding to the user's query text in a preset answer content library, and the relevance score between the answer text and the user's query text is calculated according to a preset relevance scoring formula;

[0159] The initial answer text set is obtained by initially screening multiple answer text sets based on the relevance score. The ranking rules are then constructed using the reference lexicon, and the initial answer text set is further screened based on the ranking rules to obtain the standard answer text set.

[0160] Specifically, the specific implementation method of the processor 10 for the above instructions can be referred to the description of the relevant steps in the corresponding embodiment of the accompanying drawings, and will not be repeated here.

[0161] Furthermore, if the modules / units integrated in the electronic device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a storage medium. The storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).

[0162] The present invention also provides a storage medium storing a computer program, which, when executed by a processor of an electronic device, can perform the following:

[0163] Obtain multiple scenario terms in the target application scenario, search for scenario keywords in the multiple scenario terms, and perform relationship analysis on the multiple scenario terms to obtain the word relationships;

[0164] A reference lexicon is generated based on multiple scenario keywords and the word relationships. The reference lexicon is then used to expand the preset information search engine to obtain a standard search engine.

[0165] The standard search engine is used to search for one or more answer texts corresponding to the user's query text in a preset answer content library, and the relevance score between the answer text and the user's query text is calculated according to a preset relevance scoring formula;

[0166] The initial answer text set is obtained by initially screening multiple answer text sets based on the relevance score. The ranking rules are then constructed using the reference lexicon, and the initial answer text set is further screened based on the ranking rules to obtain the standard answer text set.

[0167] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0168] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0169] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0170] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0171] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.

[0172] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0173] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0174] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in the system claims may also be implemented by a single unit or device through software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any specific order.

[0175] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for information query based on section classification, characterized in that, The method comprises: Obtaining a plurality of scene vocabularies in a target application scenario, searching for scene keywords in the plurality of scene vocabularies, and performing relationship analysis on the plurality of scene vocabularies to obtain vocabulary relationships; Based on the plurality of scene keywords and the vocabulary relationships, a to-be-referenced vocabulary library is constructed, and the to-be-referenced vocabulary library is used to expand a preset information search engine to obtain a standard search engine; Using the standard search engine to search in a preset answer content library to obtain one or more answer texts corresponding to a user query text, and calculating a relevance score between the answer text and the user query text according to a preset relevance scoring formula; According to the relevance score, the plurality of answer texts are initially screened to obtain an initial answer text set, the to-be-referenced vocabulary library is used to construct a grading rule, and the initial answer text set is re-screened according to the grading rule to obtain a standard answer text set; Wherein, the relationship analysis on the plurality of scene vocabularies to obtain the vocabulary relationships comprises: identifying the vocabulary types corresponding to the scene vocabularies, and constructing a hierarchical relationship or a synonymous relationship between the scene vocabularies corresponding to the same vocabulary type; according to a preset corresponding rule, the scene vocabularies corresponding to different vocabulary types are constructed to have a corresponding relationship; and the hierarchical relationship, the synonymous relationship and the corresponding relationship are summarized as the vocabulary relationships; The initial screening of the plurality of answer texts according to the relevance score to obtain the initial answer text set comprises: sorting the answer texts in the plurality of answer texts in descending order of the relevance score to obtain a text ranking list; and selecting a plurality of answer texts in the front of the text ranking list as the initial answer text set; The construction of the grading rule using the to-be-referenced vocabulary library comprises: analyzing the vocabularies in the to-be-referenced vocabulary library, taking the scene keywords in the analyzed vocabularies as first-level vocabularies, and setting a matching rule for successful matching with the first-level vocabularies; performing vocabulary statistics on the vocabularies that meet the initial rule to obtain a number of vocabularies, constructing an initial grading rule according to the interval in which the number of vocabularies is located; and summarizing the matching rule and the initial grading rule to obtain the grading rule.

2. The method for information search based on section classification according to claim 1, wherein, The searching for the scene keywords in the plurality of scene vocabularies comprises: Extracting a plurality of scene vocabularies from preset historical scene data, and summarizing the plurality of scene vocabularies into a scene training set; Using the scene training set to train a preset convolutional network model to obtain a keyword extraction model; According to the keyword extraction model, the scene vocabularies are subjected to keyword extraction to obtain scene keywords.

3. The method for information search based on section classification according to claim 1, wherein, The re-screening of the initial answer text set according to the grading rule to obtain the standard answer text set comprises: Extracting initial answer texts with consistent relevance scores in the initial answer text set as a target text set; Grading target texts in the target text set based on the grading rule to obtain a standard answer text set.

4. The method for information search based on section classification according to claim 1, wherein, The calculation of the relevance score between the answer text and the user query text according to the preset relevance scoring formula comprises: The preset relevance scoring formula is: in, The relevance score is... For the user's query text, The answer text, For the first in the answer text One keyword, The inverse text frequency of the keywords in the user query text. and These are the preset adjustment parameters. This is the preset document average value. The modulus corresponding to the answer text. Indicates the first Key words In the answer text The frequency of occurrence.

5. The method for information search based on section classification according to claim 1, wherein, The standard search engine is obtained by extending the preset information search engine by using the to-be-referenced vocabulary. The information search engine is used to identify information of the to-be-referenced vocabulary, and a to-be-added word is obtained. The to-be-added word is added to an extended vocabulary of a segmenter in the information search engine, and the standard search engine is obtained.

6. An information search device based on section classification, for implementing the information search method based on section classification according to any one of claims 1 to 5, characterized by The device comprises: The relationship analysis module is configured to acquire a plurality of scene words in a target application scenario, search for scene keywords in the plurality of scene words, and analyze relationships of the plurality of scene words to obtain word relationships. The engine extension module is configured to construct a to-be-referenced vocabulary based on the scene keywords and the word relationships, extend a preset information search engine by using the to-be-referenced vocabulary, and obtain a standard search engine. The score calculation module is configured to search for one or more answer texts corresponding to a user query text in a preset answer content library by using the standard search engine, and calculate a relevance score between the answer texts and the user query text according to a preset relevance scoring formula. The double screening module is configured to perform a primary screening on the answer texts according to the relevance score to obtain an initial answer text set, construct a grading rule by using the to-be-referenced vocabulary, and perform a secondary screening on the initial answer text set according to the grading rule to obtain a standard answer text set.

7. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected to the at least one processor in communication; wherein The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the information query method based on the fixed section grading according to any one of claims 1 to 5.

8. A storage medium storing a computer program, characterized by The computer program is executed by the processor to implement the information query method based on the fixed section grading according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and system for network information extraction

    CN101344889A

  • Searching method and device based on keywords and semantics, equipment and storage medium

    CN115438166A