Dialogue analysis retrieval method, device and medium

By generating big model prompt words and calling big model to generate synonyms, the problem that the big model dialogue analysis system is difficult to match user problems in relational data retrieval, and efficient and accurate synonyms and business knowledge are achieved, reducing labor costs and promoting system intelligence.

CN119938890BActive Publication Date: 2025-06-13INSPUR GENERSOFT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510429242.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-06-13
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The existing large-model dialogue analysis system is difficult to accurately match the fields and dictionary items in user colloquialization problems in relational data retrieval, and the user's problems do not clearly indicate business knowledge, which makes the knowledge unable to be known by the big model, and manually configuring business knowledge is labor-intensive and inefficient.

Method used

By obtaining data structure information, generating big model prompt words, calling big model to generate fields and dictionary synonyms, performing knowledge search and association, and updating the knowledge base for analysis and search.

Benefits of technology

It reduces manual intervention, reduces labor cost investment, improves the accuracy of synonyms and the efficiency of obtaining business knowledge, and promotes the intelligent development of dialogue analysis systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938890B_ABST
    Figure CN119938890B_ABST
Patent Text Reader

Abstract

The present application discloses a conversational analysis and retrieval method, device, and medium, which relate to the field of data processing. The method includes: obtaining corresponding data structure information for a preset data set; generating a first large model prompt word corresponding to the original content through a pre-set synonym generation template; calling a large model based on the first large model prompt word to generate synonym content; retrieving the synonym content in a knowledge base to establish an association relationship between business knowledge and the original content; and updating the knowledge base based on the association relationship. Compared with the traditional method of manually adding synonyms and knowledge, it is not necessary to consume a large amount of human and time costs. The addition of synonyms and the acquisition of knowledge can be achieved through a large model, reducing manual intervention and lowering the input of labor costs. While increasing work efficiency, it can also reduce costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and specifically to a conversational analysis retrieval method, device, and medium. Background Art

[0002] With the development of technology, large models have gradually come into view.

[0003] In traditional solutions, in a conversational analysis system based on a large model, the following key technical problems need to be solved urgently for relational data retrieval:

[0004] 1. In existing datasets, the original names in the field and dimension dictionaries are mostly configured Chinese names, while user questions are often presented in an oral form. This makes it difficult to accurately match the correct fields and the values of the enumerated items in the dictionary during Retrieval-augmented Generation (RAG) retrieval, and thus it is difficult to put the necessary knowledge into the prompt and send it to the large model.

[0005] 2. Some business knowledge may not be clearly indicated in the user's question, so that this knowledge cannot be known to the large model. Although it is possible to configure knowledge in advance on dataset fields or some dictionary items, so that it can be automatically brought out after recall. However, the input of business knowledge associated with these fields is labor-intensive, not only inefficient, but also prone to omissions, and it is difficult to meet the requirements of large-scale data processing and real-time performance. Summary of the Invention

[0006] To solve the above problems, this application proposes a conversational analysis retrieval method, including:

[0007] For a preset dataset, obtain the corresponding data structure information;

[0008] For the specified fields and / or specified dictionaries in the data structure information, as the original content, and generate the first large model prompt corresponding to the original content through a pre-set synonym generation template;

[0009] Based on the first large model prompt, call the large model to generate the corresponding field synonyms and / or dictionary synonyms as the synonym content;

[0010] Retrieve the corresponding business knowledge in the knowledge base for the synonym content, and establish an association relationship between the business knowledge and the original content;

[0011] Based on the association relationship, update the knowledge base and perform analysis and retrieval based on the updated knowledge base.

[0012] In one example, the first large model prompt corresponding to the original content is generated through a preset synonym generation template, specifically including:

[0013] For the specified field, determine its corresponding field synonym generation template; the field synonym generation template includes: field name constraint, dataset business purpose constraint, field type constraint, output format constraint;

[0014] For the specified dictionary, determine its corresponding dictionary synonym generation template; the dictionary synonym generation template includes: dictionary enumeration value constraint, affiliated field constraint, description length constraint, superior-subordinate relationship constraint, output format constraint;

[0015] According to the field synonym generation template and the dictionary synonym generation template, generate the corresponding first large model prompts for the specified field and the specified dictionary respectively.

[0016] In one example, after generating the corresponding field synonyms and / or dictionary synonyms as the synonym content by calling the large model based on the first large model prompt, the method further includes:

[0017] Determine the superior-subordinate relationship corresponding to the specified dictionary according to the data structure information;

[0018] Check for duplicates based on the dictionary synonyms to determine whether there are the same dictionary synonyms for different specified dictionaries;

[0019] If so, regard this dictionary synonym as a conflicting synonym;

[0020] Deduplicate the conflicting synonyms according to the superior-subordinate relationship of the specified dictionaries corresponding to the conflicting synonyms.

[0021] In one example, deducing the conflicting synonyms according to the superior-subordinate relationship of the specified dictionaries corresponding to the conflicting synonyms specifically includes:

[0022] Determine the multiple specified dictionaries corresponding to the conflicting synonyms;

[0023] If there is a superior-subordinate relationship among the multiple specified dictionaries, retain the dictionary synonym corresponding to the topmost specified dictionary;

[0024] If the multiple specified dictionaries are at the same level, retain the dictionary synonym corresponding to the earliest specified dictionary according to the order of the specified dictionaries.

[0025] In one example, after determining the multiple specified dictionaries corresponding to the conflicting synonyms, the method further includes:

[0026] Determine whether there is a preset core term in the multiple specified dictionaries or the conflicting synonyms;

[0027] If there is a core term in the multiple specified dictionaries, ignore the hierarchical relationship and retain the dictionary synonyms corresponding to the core term;

[0028] If the conflicting synonym is a core term, determine a template according to the preset approximation degree to generate a second largest model prompt word corresponding to the core term; the approximation degree determination template includes: dictionary enumeration value constraint, belonging field constraint, core term constraint, output format constraint;

[0029] Based on the second largest model prompt word, call a large model to determine the similarity between the core term and each specified dictionary, and determine the specified dictionary corresponding to the retained core term according to the similarity and the hierarchical relationship corresponding to the multiple specified dictionaries.

[0030] In one example, after generating corresponding field synonyms and / or dictionary synonyms as synonym content by calling a large model based on the first largest model prompt word, the method further includes:

[0031] Determine the specified dictionary and the corresponding dictionary synonyms, and determine the business scenario to which it belongs according to the specified dictionary;

[0032] If there is a hit on the semantic disabling relationship corresponding to the preset business scenario for the specified dictionary and the dictionary synonyms, delete the hit dictionary synonyms;

[0033] If the business scenario belongs to a preset scenario, generate a third largest model prompt word corresponding to the dictionary synonyms according to the preset synonym screening template; the synonym screening template includes: dictionary enumeration value constraint, belonging field constraint, business scenario constraint, output format constraint;

[0034] Based on the third largest model prompt word, call a large model to screen the dictionary synonyms.

[0035] In one example, the analysis and retrieval based on the updated knowledge base specifically includes:

[0036] Obtain the retrieval content input by the user;

[0037] Perform word segmentation on the retrieval content to obtain a number of entity words;

[0038] According to the entity words, recall the entity information corresponding to relevant dictionaries and relevant fields in the knowledge base;

[0039] According to the entity information, determine the associated knowledge with an association relationship in the knowledge base;

[0040] Convert the entity information and the associated knowledge into structured language, and perform analysis and retrieval through the structured language.

[0041] In one example, converting the entity information and the associated knowledge into structured language specifically includes:

[0042] Determine the dimensional attributes of the entity information;

[0043] According to the dimensional attributes, determine whether the entity information corresponds to a specified field or a specified dictionary;

[0044] If it corresponds to a specified field, determine that the entity information is the first weight;

[0045] If it corresponds to a specified dictionary, determine whether the entity information belongs to the original content;

[0046] If it belongs to the original content, determine that the entity information is the second weight;

[0047] If it does not belong to the original content, determine whether the entity information belongs to the synonym content temporarily obtained this time;

[0048] If it belongs to the temporarily obtained synonym content, determine that the entity information is the third weight;

[0049] If it does not belong to the temporarily obtained synonym content, determine the update time of the entity information in the knowledge base, and determine the fourth weight corresponding to the entity information according to the update time;

[0050] According to the weight corresponding to the entity information, convert the entity information and the associated knowledge into structured language;

[0051] Among them, the weight values are from high to low in turn: the first weight, the second weight, the third weight, and the fourth weight.

[0052] On the other hand, the present application also proposes a conversational analysis and retrieval device, including:

[0053] At least one processor; and,

[0054] A memory communicatively connected to the at least one processor; wherein,

[0055] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the conversational analysis and retrieval method as described in any of the above examples.

[0056] On the other hand, the present application also proposes a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are configured as: the conversational analysis and retrieval method described in any of the above examples.

[0057] The conversational analysis and retrieval method proposed by the present application can bring the following beneficial effects:

[0058] 1. Compared with the traditional method of manually adding synonyms and knowledge, there is no need to consume a large amount of human and time costs. The addition of synonyms and the acquisition of knowledge can be achieved through large models, reducing manual intervention, lowering the input of labor costs, increasing work efficiency, and reducing costs at the same time.

[0059] 2. Compared with traditional manual addition, using large model technology can generate more accurate synonyms that are more in line with the business scenario, thereby increasing the accuracy of the finally returned knowledge.

[0060] 3. By deeply integrating large model technology into the relational data retrieval process, it promotes the development of the conversational analysis system towards the intelligent direction. The function of automatically generating synonyms enables the system to better simulate the human language understanding method, interact with users more naturally, improves the intelligent level and user experience of the entire system, and lays a foundation for the intelligent upgrade of related fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:

[0062] Figure 1 is a schematic flowchart of the conversational analysis and retrieval method in an embodiment of the present application;

[0063] Figure 2 is a schematic diagram of the conversational analysis and retrieval method in a certain situation in an embodiment of the present application;

[0064] Figure 3 is a schematic diagram of the conversational analysis and retrieval device in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0066] The following will, in conjunction with the accompanying drawings, elaborate in detail on the technical solutions provided by each embodiment of the present application.

[0067] To address the above-mentioned issues, an algorithm for generating synonyms can be constructed around probability models and calculation logics. For example, a statistics-based algorithm can be set up to determine synonyms by analyzing the co-occurrence frequencies of words in a large-scale text corpus. The core lies in judging which words have similar semantics based on the probability relationships of words appearing in the text. There is also a deep learning-based algorithm that learns the vector representations of words with the help of a neural network model and identifies synonyms by calculating the similarity between vectors.

[0068] However, in the data analysis scenario, the situation may become more complex. Although they are all synonyms, there may be differences among the dataset field context, the dimension table dictionary context, and the business knowledge context. Taking the dataset field context as an example, each field has its specific data meaning and usage. For instance, in financial data analysis, the "sales amount" field represents the revenue amount obtained by an enterprise through selling products or services. Synonyms in this context need to closely revolve around this business concept, and the knowledge of personalized differences of these concepts in a specific company also needs to be fed into the large model. The dimension table dictionary context focuses on the definition and interpretation of data dimensions. For example, in a dimension table dictionary containing the region dimension, "province" may correspond to the specific name of a provincial administrative region, and the generation of its synonyms needs to fit this dimension definition context. The business knowledge context is more macroscopic, covering aspects such as the professional knowledge, industry rules, and business processes of the entire business field. For example, in the medical data analysis scenario, the business knowledge context includes information such as medical terms, disease diagnosis criteria, and treatment processes, and the generation of synonyms must conform to these professional knowledge backgrounds.

[0069] Traditional synonym generation algorithms often struggle to make full use of this additional information. Since most of these algorithms are trained based on general natural language texts and are not optimized for the special contexts in the data analysis scenario, when dealing with the generation of synonyms related to dataset fields, dimension table dictionaries, and business knowledge, the generated synonyms may not be accurate or may not meet the actual business requirements. For example, in general texts, "spending" and "expenditure" may be regarded as synonyms, but in the context of financial data analysis, "spending" may be more focused on the concept of daily consumption, while "expenditure" is more financially professional and suitable for formal financial statement records. Existing general algorithms may have difficulty accurately distinguishing such context differences.

[0070] Based on this, a solution for private domain knowledge can establish a knowledge base for the private domain. During operation, relevant knowledge is first recalled from the knowledge base, and then precise ranking or reasoning is performed to eliminate interference items. For example, when a user's question mentions an entity, such as how a certain indicator is this year. During recall, only the indicator information of that indicator can be recalled based on the question, and the knowledge related to that indicator needs to be recalled again in the knowledge base. Since the knowledge base retrieves based on a similarity algorithm, this process will increase the probability of making mistakes. Obviously, the second recall can solve this problem by storing the indicator ID and related knowledge in a structured manner. However, it is necessary to associate the indicator ID with the relevant knowledge in advance, usually relying on manual work, which is time-consuming and laborious.

[0071] Based on this, as Figure 1 shown, the embodiments of the present application provide a conversational analysis retrieval method to solve the problems existing in the above traditional solutions.

[0072] As Figure 1 and Figure 2 shown, the method includes:

[0073] S101: For a preset data set, obtain the corresponding data structure information.

[0074] A conversational analysis retrieval system (hereinafter simply referred to as the analysis system) means that when a user asks a question to the analysis system, the analysis system retrieves corresponding knowledge according to the question asked by the user and then gives a corresponding answer.

[0075] The analysis system generally writes one or more complete query SQLs for the theme to be analyzed to form a queryable data view, and this data view can be called a data set.

[0076] The analysis system is usually associated with a Business Intelligence System (BI system) or is a system module in the BI system. The BI system can automatically identify the data field types, which is convenient for the analysis system to perform corresponding analysis and retrieval.

[0077] And implementers usually manually mark the name of each field, manually give a name to this data set that can represent the meaning of the data set, add a data set description to the data set, and make remarks on the included data.

[0078] Based on this, traverse the dataset by theme to obtain the corresponding data structure information. The data structure information may include: a brief description of the theme, the dataset name, the dataset description, the name of each field in the dataset, the field type information, and the dimension dictionary information associated with the field. Read the text description, level, and superior-subordinate relationship of each enumeration item in each dimension dictionary according to the dimension dictionary information. Among them, the superior-subordinate relationship refers to a semantic relationship of inclusion and being included. For example, it includes geographical inclusion such as Province A and City B within Province A, and semantic inclusion such as total sales and sales of a certain product. Of course, the superior-subordinate relationship here mainly refers to the superior-subordinate relationship within the specified dictionary. If a specified dictionary has a superior relationship, but this superior relationship does not belong to the specified dictionary, it does not need to be considered.

[0079] S102: For the specified field and / or specified dictionary in the data structure information, use it as the original content, and generate the first large model prompt word corresponding to the original content through a pre-set synonym generation template.

[0080] The specified field and specified dictionary refer to some or all of the original fields and dictionaries in the data table selected according to requirements or settings. Among them, the dictionary can also be called the dimension dictionary or dimension table dictionary, which refers to the specific value of a certain piece of data under the field.

[0081] Specifically, for the specified field, determine its corresponding field synonym generation template (which can also be called the field synonym Prompt template); the field synonym generation template includes: field name constraint, dataset business purpose constraint, field type constraint, and output format constraint.

[0082] Generally speaking, the field synonym generation templates for each field are the same, but different templates can also be set based on requirements.

[0083] For example, the field synonym generation template can be: <! Field name!> is a column in the data table <! Dataset name!>, the business purpose of the table is <! Dataset description!>, the field type is <! Field type!>, please generate possible synonyms for the field and return them in the form of a JSON array.

[0084] Among them, <! String!> means the system variable corresponding to this "string", which can be replaced with the content in the data structure information. At this time, the variable corresponding to the field name corresponds to the field name constraint, the variable corresponding to the dataset description corresponds to the dataset business purpose constraint, and the variable corresponding to the field type corresponds to the field type constraint. Returning in the form of a JSON array corresponds to the output format constraint.

[0085] Similarly, for a specified dictionary, determine its corresponding dictionary synonym generation template; the dictionary synonym generation template includes: dictionary enumeration value constraints, affiliated field constraints, description length constraints, superior-subordinate relationship constraints, and output format constraints.

[0086] For example, the dictionary synonym generation template can be: <!dictionary enumeration value!> is one of the values in the column <!field name!> of the data table <!dataset name!>, the field length is <!field length!>, its superior is <!superior name!>, and it has <!number of subordinates!> subordinates. Please generate possible synonyms and return them in the form of a JSON array.

[0087] Among them, the dictionary enumeration value refers to the specific value corresponding to each specified dictionary under each field. The variable corresponding to the dictionary enumeration value corresponds to the dictionary enumeration value constraint, the variable corresponding to the field name corresponds to the affiliated field constraint, the variable corresponding to the field length (here it refers to the length of this dictionary) corresponds to the description length constraint, the variables corresponding to the superior name and the number of subordinates correspond to the superior-subordinate relationship constraint, and returning in the form of a JSON array corresponds to the output format constraint.

[0088] Based on this, according to the field synonym generation template and the dictionary synonym generation template, generate corresponding first large model prompt words for the specified field and the specified dictionary respectively.

[0089] For example, the first large model prompt word corresponding to the specified field can be: The product name is a column in the data table e-commerce sales situation table. The business purpose of the table is to describe the sales situation such as the sales amount and sales volume of each product on the e-commerce platform within a certain period of time. The field type is varchar. Please generate possible synonyms for the field and return them in the form of a JSON array.

[0090] The first large model prompt word corresponding to the specified dictionary can be: Tablet computer is one of the values in the column product name of the data table e-commerce sales situation table. The field length is varchar(10), its superior is computer, and it has 10 subordinates. Please generate possible synonyms and return them in the form of a JSON array.

[0091] S103: Based on the first large model prompt word, call the large model to generate corresponding field synonyms and / or dictionary synonyms as the synonym content.

[0092] At this time, for each first large model prompt word, call the large model API to generate and return the corresponding synonym content through the large model. For the same original content, it can include one or more synonym contents. In some scenarios, it can also be set to not find the corresponding synonym content.

[0093] The obtained field synonyms can be recorded in the field information corresponding to the specified field, and the obtained dictionary synonyms can also be recorded in the corresponding dictionary name structure.

[0094] Furthermore, after obtaining the synonym content, since the specified field and the specified dictionary both come from the same data table, there is likely to be a situation where the same synonym content exists. For example, there are two specified dictionaries, namely "tablet computer" and "desktop computer", and they may both have the same synonym "computer".

[0095] At this time, if the synonym content is directly used, it is easy to cause content duplication, which not only wastes resources but also easily causes subsequent content conflicts. Therefore, duplicate removal processing can be performed after obtaining the synonym content.

[0096] According to the data structure information, determine the hierarchical relationship corresponding to the specified dictionary. As mentioned above, the hierarchical relationship corresponding to each specified dictionary has been determined in the data structure information. In the hierarchical relationship, multiple levels can be set in advance, and the level where each specified dictionary is located can be set for each specified dictionary, so as to obtain the corresponding hierarchical relationship.

[0097] For example, three levels are set. The first level is the highest level, which includes two specified dictionaries, namely "Region A" and "Computer", and the second level is the medium level, which includes "Area B", "Area C", "Tablet Computer", "Desktop Computer", etc. Among them, "Area B" and "Area C" both belong to "Region A". At this time, "Area B" and "Area C" are in a peer relationship, and their relationship with "Region A" is a subordinate relationship. Similarly, "Tablet Computer" and "Desktop Computer" are also in a peer relationship, and their relationship with "Computer" is a subordinate relationship. Of course, to ensure the clarity of the hierarchical relationship, corresponding connection relationships can be set to clarify which specified dictionaries have hierarchical relationships between different levels.

[0098] Perform duplicate checking based on the dictionary synonyms to determine whether there are the same dictionary synonyms in different specified dictionaries.

[0099] If not, there is no need to perform duplicate removal processing.

[0100] If there are, then regard the dictionary synonym as a conflicting synonym. At this time, according to the hierarchical relationship of the specified dictionary corresponding to the conflicting synonym, perform duplicate removal on the conflicting synonym.

[0101] Specifically, when performing duplicate removal, determine multiple specified dictionaries corresponding to the conflicting synonym.

[0102] If there is a hierarchical relationship among multiple specified dictionaries, retain the dictionary synonyms corresponding to the topmost specified dictionary. That is, use this conflicting synonym as the dictionary synonym corresponding to the topmost specified dictionary, and delete this conflicting synonym corresponding to other specified dictionaries.

[0103] If multiple specified dictionaries are at the same level, retain the dictionary synonyms corresponding to the earliest specified dictionary according to the order of the specified dictionaries.

[0104] Generally speaking, only for specified dictionaries with a hierarchical relationship (including parent-child relationship, sibling relationship), their synonyms may be the same. Of course, in extremely rare cases, there may also be conflicting synonyms between two dictionaries without a hierarchical relationship. In this case, the hierarchical relationship between the two can be judged according to the level where the dictionary is located.

[0105] For specified dictionaries with a parent-child relationship or a sibling relationship, retaining the dictionary synonyms corresponding to the topmost specified dictionary can store this conflicting synonym later to prevent the lower level from reusing the name of the higher level, ensure the integrity of category data, and prevent contamination by redundant data. For example, assume that in the specified dictionary, "Product D" is the parent of "Product E", and their corresponding conflicting synonym is "Product F". At this time, there may be three situations for Product F, namely: First, Product F is at a higher level than Product D; second, Product F is at the same level as Product D; third, Product F is at the same level as or lower than Product E.

[0106] At this time, for the first situation, whether retaining this conflicting synonym "Product F" as the synonym of "Product D" or as the synonym of "Product E", when the user searches for Product D or Product E, products that do not belong to themselves will appear, resulting in data contamination. However, at this time, since the level of Product D is higher and its scope is larger, when retaining this conflicting synonym "Product F" as the synonym of "Product D", the resulting data contamination is relatively less.

[0107] Of course, if the levels in the hierarchical relationship are reasonably divided in advance, the first situation usually does not occur.

[0108] For the second situation, when retaining this conflicting synonym as the synonym of "Product D", since they are at the same level, there will be no contamination when the user searches for Product D. When retaining this conflicting synonym as the synonym of "Product E", since the level of Product F is higher than the level of Product E, there will still be data contamination when the user searches for Product E, and products that belong to Product F but not to Product E will appear.

[0109] For the third situation, there will be no data contamination in both cases.

[0110] Based on this, retaining the dictionary synonyms corresponding to the top-level specified dictionary can increase the subsequent search scope of users while reducing the occurrence of data pollution. For the specified dictionaries at the same level, the effects of retaining the dictionary synonyms for any one of the specified dictionaries are basically similar, so they can be randomly selected, or retained in the order of appearance.

[0111] S104: Retrieve in the knowledge base for the synonym content, recall the corresponding business knowledge, and establish an association relationship between the business knowledge and the original content.

[0112] Through the business rules stored in the knowledge base, the business knowledge corresponding to the synonym content can be retrieved, so as to obtain the required business knowledge. For each synonym content, retrieve it separately, and then establish an association relationship between the retrieved business knowledge and the original content corresponding to the synonym content.

[0113] The association relationship here usually refers to the association relationship stored in the cache. Through this association relationship, when the user retrieves through the original content currently, the business knowledge associated with the corresponding synonym content can be retrieved.

[0114] S105: Update the knowledge base based on the association relationship, and perform analysis and retrieval based on the updated knowledge base.

[0115] In order to ensure that in the subsequent process, the business knowledge corresponding to the synonym content can still be retrieved through this association relationship, the knowledge base can be updated according to this association relationship, and this association relationship is written into the knowledge base, so that in the subsequent retrieval process of the user, relevant business knowledge can also be retrieved.

[0116] Among them, the update can be carried out in a regular update manner. For example, the previously obtained association relationship is temporarily stored in the cache. After a certain period of time, the association relationship in the cache is updated to the knowledge base, and the corresponding update timestamp is recorded.

[0117] 1. Compared with the traditional way of manually adding synonyms and knowledge, there is no need to consume a large amount of human and time costs. Through the large model, the addition of synonyms and the acquisition of knowledge can be realized, reducing manual intervention, lowering the input of human costs, increasing work efficiency, and reducing costs at the same time.

[0118] 2. Compared with traditional manual addition, using large model technology can generate more accurate synonyms that are more in line with the business scenario, thereby increasing the accuracy of the finally returned knowledge.

[0119] 3. By deeply integrating big model technology into the relational data retrieval process, the conversational analysis system has been promoted to develop in an intelligent direction. The function of automatically generating synonyms allows the system to better simulate the way human language is understood and interact with users more naturally, which improves the intelligence level and user experience of the entire system and lays the foundation for the intelligent upgrade of related fields.

[0120] In one embodiment, when performing deduplication of conflicting synonyms, there may be some special cases, so that the final effect is not satisfactory if only the deduplication method described above is used.

[0121] Specifically, for some special commodities, or terms in special industries (for example, the medical industry), there are some core words, which may not be at a high level, but are relatively more important.

[0122] At this time, it is determined whether there are preset core terms in multiple designated dictionaries or conflicting synonyms. The core terms can be obtained from relevant industry dictionaries, operation manuals, product brochures, expert experience, etc.

[0123] If a core term exists in multiple specified dictionaries, the hierarchical relationship is ignored and the dictionary synonyms corresponding to the core term are retained. Even if the core term has a lower level, the hierarchical relationship is no longer considered, the dictionary synonyms corresponding to the core term are retained, and the dictionary synonyms corresponding to other specified dictionaries are deleted, thereby highlighting the importance of the core term.

[0124] If the conflicting synonyms are core terms, considering the particularity of the core terms, only considering the hierarchical relationship may cause confusion in the synonym relationships that are ultimately retained.

[0125] Therefore, according to the preset proximity determination template, the second largest model prompt word corresponding to the core term is generated; the proximity determination template includes: dictionary enumeration value constraints, belonging field constraints, core term constraints, and output format constraints.

[0126] For example, the proximity determination template may be:<!字典枚举值!> and<!字典枚举值!> All are data tables<!数据集名称!> Please determine whether it matches a value in a column <! Field name! ><!核心术语枚举值!> The similarity between the two is output as a percentage and returned as a JSON array.

[0127] Among them, the similarity determination template can be set to multiple, and the<!字典枚举值!> The number of is different, corresponding to the number of specified dictionaries.

[0128] Based on the second largest model prompt word, the large model is called to determine the similarity between the core term and each specified dictionary.

[0129] Based on the similarity and the hierarchical relationships corresponding to multiple specified dictionaries, determine the specified dictionary corresponding to the core term to be retained. For example, set corresponding weights for the similarity and hierarchical relationships, obtain the corresponding score through weighted summation, and select the specified dictionary with a higher score, retaining the core term as a dictionary synonym. Among them, the score corresponding to the hierarchical relationship can be set to correspond to a score for each level difference, so as to obtain the score corresponding to the hierarchical relationship according to the level of each specified dictionary, and perform weighted summation with the score corresponding to the similarity to obtain the final score.

[0130] In one embodiment, in addition to the special cases corresponding to the core terms mentioned above, there may be some other special cases that affect the deduplication effect.

[0131] Specifically, for the special requirements of some merchants, or the special requirements of the industry, in order to ensure the particularity of their own products, or the industry requires the independence of certain terms, it is prohibited that there is an associated relationship between certain words and some common synonyms.

[0132] Based on this, determine the specified dictionary and the corresponding dictionary synonyms, and determine the business scenario to which they belong according to the specified dictionary. The business scenario can be obtained by whether the specified dictionary hits relevant keywords. For example, pre-select some common keywords and establish relationships with each business scenario.

[0133] For each business scenario, corresponding semantic disabling relationships can be set according to relevant industry dictionaries, operation manuals, product brochures, expert experience, etc., which contain some disabling relationships that prohibit which words from being synonyms. And setting different business scenarios with different semantic disabling relationships can reduce the complexity of comparison and duplicate checking and increase efficiency in the case of a large number of semantic disabling relationships.

[0134] At this time, if there are a specified dictionary and dictionary synonyms that hit the semantic disabling relationship corresponding to the pre-set business scenario, then delete the hit dictionary synonym. The hit here requires that the specified dictionary and dictionary synonyms hit both words in a single semantic disabling relationship at the same time. If only one of the specified dictionary and dictionary synonyms hits one word in a certain semantic disabling relationship, it is not considered a hit of that semantic disabling relationship.

[0135] If the business scenario belongs to a pre-set scenario (such as industries with strict requirements like the medical industry, legal industry, etc.), then in addition to setting this semantic disabling relationship, a third large model prompt word corresponding to the dictionary synonym can also be generated according to the pre-set synonym screening template; the synonym screening template includes: dictionary enumeration value constraint, belonging field constraint, business scenario constraint, output format constraint.

[0136] For example, the synonym screening template can be: <!Dictionary enumeration value!> is one of the values in a column <!Field name!> in the data table <!Dataset name!>, and it has the following synonyms: <!Synonym enumeration value!>, <!Synonym enumeration value!>. The business scenario of this <!Dictionary enumeration value!> is <!Business scenario name!>. Please screen out the above synonyms and only retain the synonyms that match this business scenario. And return it in the form of a JSON array.

[0137] Among them, multiple synonym screening templates can be set, and the number of <!Synonym enumeration value!> in each template is different, which is used to correspond to the number of dictionary synonyms.

[0138] Based on the third large model prompt, call the large model to screen the dictionary synonyms, so as to ensure the rationality of synonyms in some important industries and businesses.

[0139] In one embodiment, when the user performs analysis and retrieval through the updated knowledge base, the retrieval content input by the user is obtained. The retrieval content is segmented to obtain several entity words. The segmentation process can be achieved by rule-based methods such as forward maximum matching and reverse maximum matching, or by traditional probability models such as hidden Markov models and conditional random fields, or by deep learning models such as BERT and BiLSTM-CRF.

[0140] According to the entity words, recall the entity information corresponding to the relevant dictionaries and relevant fields in the knowledge base. This entity information can include the entity information corresponding to the fields and dictionaries.

[0141] According to the entity information, determine the associated knowledge with an associated relationship in the knowledge base. At this time, this associated knowledge can include the original associated knowledge corresponding to the entity words, or the business knowledge with an associated relationship brought by synonyms after the knowledge base is updated.

[0142] According to the entity information and the associated knowledge, convert it into a structured language and perform analysis and retrieval through the structured language. At this time, it can be input into the NL2SQL agent to improve the accuracy. Among them, the NL2SQL agent is an intelligent system based on artificial intelligence technology, and its core function is to automatically convert the natural language instructions input by the user into executable SQL statements, so as to achieve efficient query and management of the database.

[0143] Furthermore, when converting structured language, due to the business knowledge in the knowledge base, which may include business knowledge directly obtained from the original content and business knowledge obtained from synonym content in different time periods respectively, the amount of business knowledge data is large. If it is directly used, it may lead to chaos in business knowledge, and ultimately the retrieved results returned to the user do not fully meet the user's needs.

[0144] Based on this, determine the dimensional attributes of the entity information. Its dimensional attributes may include: belonging to a field or a dictionary, whether it belongs to a synonym, whether it belongs to a synonym obtained temporarily this time, the update time of this synonym in the knowledge base, etc. For example, for a certain entity information, if it is a specified dictionary newly added by the user in this search, its dimensional attributes are: [dictionary; belonging; belonging; empty]. For the convenience of description, this dimensional attribute can be described by corresponding codes. For example, set the code for "belonging" as 1, the code for "not belonging" as 0, the code for belonging to a field as 1, the code for belonging to a dictionary as 0, etc.

[0145] According to the dimensional attributes, determine whether the entity information corresponds to a specified field or a specified dictionary. If it corresponds to a specified field, determine that the entity information has the first weight. Compared with a dictionary, the content contained in a field is more general and can highlight the essential purpose of the user's question. Therefore, it is judged first, and when it is a field, it can be considered that the highest first weight is set for it to represent the user's core purpose.

[0146] If it corresponds to a specified dictionary, continue to judge and determine whether the entity information belongs to the original content. If it belongs to the original content, determine that the entity information has the second weight. Compared with the business knowledge obtained from synonym content, the original content is more in line with the user's initial idea. Therefore, if it belongs to the original content of the specified dictionary, a second weight lower than the first weight is set for it.

[0147] If it does not belong to the original content, continue to judge and determine whether the entity information belongs to the synonym content obtained temporarily this time. If it belongs to the synonym content obtained temporarily, determine that the entity information has the third weight. The synonym content obtained temporarily is more in line with the current scenario requirements and can better ensure the timeliness of synonyms compared with the content that has been updated in the knowledge base before. Therefore, if it belongs to the synonym content obtained temporarily this time in the specified dictionary, a third weight lower than the second weight is set for it.

[0148] If it does not belong to the temporarily obtained synonym content, determine the update time of the entity information in the knowledge base, and determine the fourth weight corresponding to the entity information according to the update time. Similarly, if it belongs to the synonym content that has been updated in the knowledge base, the fourth weight lower than the third weight can be set according to the update time. Among them, the closer the update time is to the current time, the higher its fourth weight, so as to ensure the timeliness of the entity information.

[0149] At this time, according to the weight corresponding to the entity information, convert the entity information and associated knowledge into structured language, so as to ensure that the finally returned data meets the needs of users.

[0150] Such as Figure 3 shown, the embodiment of the present application also provides a conversational analysis retrieval device, including:

[0151] At least one processor; and,

[0152] A memory communicatively connected to the at least one processor; wherein,

[0153] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the conversational analysis retrieval method as described in any of the above examples.

[0154] The embodiment of the present application also provides a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set to: the conversational analysis retrieval method as described in any of the above examples.

[0155] Each embodiment in the present application is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

[0156] The device and medium provided by the embodiment of the present application correspond one-to-one with the method. Therefore, the device and medium also have beneficial technical effects similar to those of the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the device and medium will not be elaborated here.

[0157] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A conversational analysis retrieval method, characterized in that: include: For the preset data set, obtain the corresponding data structure information; For the designated field and / or designated dictionary in the data structure information, as the original content, and through a preset synonym generation template, a first large model prompt word corresponding to the original content is generated; Based on the first large model prompt word, calling the large model, generating corresponding field synonyms and / or dictionary synonyms as synonym content; Searching the synonymous content in the knowledge base, recalling the corresponding business knowledge, and establishing an association relationship between the business knowledge and the original content; Based on the association relationship, the knowledge base is updated, and analysis and retrieval are performed based on the updated knowledge base.

2. The conversational analysis retrieval method according to claim 1, characterized in that: Generate the first model prompt word corresponding to the original content by using a preset synonym generation template, specifically including: For the specified field, determine the corresponding field synonym generation template; the field synonym generation template includes: field name constraints, data set business purpose constraints, field type constraints, and output format constraints; For the specified dictionary, determine the corresponding dictionary synonym generation template; the dictionary synonym generation template includes: dictionary enumeration value constraints, belonging field constraints, description length constraints, superior-subordinate relationship constraints, and output format constraints; According to the field synonym generation template and the dictionary synonym generation template, corresponding first large model prompt words are generated for the designated field and the designated dictionary respectively.

3. The conversational analysis retrieval method according to claim 1, characterized in that: Based on the first large model prompt word, calling the large model to generate corresponding field synonyms and / or dictionary synonyms as synonym content, the method further includes: Determine, according to the data structure information, the superior-subordinate relationship corresponding to the designated dictionary; Perform a duplicate check based on the dictionary synonyms to determine whether the same dictionary synonyms exist in different designated dictionaries; If it exists, the dictionary synonym is used as the conflicting synonym; The conflicting synonyms are deduplicated according to the superior-subordinate relationship of the designated dictionaries corresponding to the conflicting synonyms.

4. The conversational analysis retrieval method according to claim 3, characterized in that: Deduplication of the conflicting synonyms is performed according to the superior-subordinate relationship of the designated dictionaries corresponding to the conflicting synonyms, specifically including: Determining a plurality of designated dictionaries corresponding to the conflicting synonyms; If there is a superior-subordinate relationship among the multiple designated dictionaries, the dictionary synonyms corresponding to the top-level designated dictionary are retained; If the multiple designated dictionaries are in a peer relationship, then according to the order of the designated dictionaries, the dictionary synonyms corresponding to the frontmost designated dictionary are retained.

5. The conversational analysis retrieval method according to claim 4, characterized in that: After determining the multiple designated dictionaries corresponding to the conflicting synonyms, the method further includes: Determining whether there is a preset core term in the plurality of designated dictionaries or the conflicting synonyms; If a core term exists in the multiple specified dictionaries, the superior-subordinate relationship is ignored, and the dictionary synonyms corresponding to the core term are retained; If the conflicting synonym is a core term, a second largest model prompt word corresponding to the core term is generated according to a preset proximity determination template; the proximity determination template includes: dictionary enumeration value constraints, belonging field constraints, core term constraints, and output format constraints; Based on the second large model prompt word, calling the large model to determine the similarity between the core term and each specified dictionary; And according to the similarity and the hierarchical relationship corresponding to the multiple designated dictionaries, the designated dictionary corresponding to the core term to be retained is determined, including: setting corresponding weights for the similarity and the hierarchical relationship, obtaining corresponding scores by weighted summation, and selecting the designated dictionary with a higher score, and retaining the core term as a dictionary synonym; wherein the score corresponding to the hierarchical relationship is set to one score for each level difference, and the score corresponding to the hierarchical relationship is obtained according to the level of the levels between the designated dictionaries.

6. The conversational analysis retrieval method according to claim 1, characterized in that: Based on the first large model prompt word, calling the large model to generate corresponding field synonyms and / or dictionary synonyms as synonym content, the method further includes: Determine the designated dictionary and the corresponding dictionary synonyms, and determine the business scenario according to the designated dictionary; If the specified dictionary and the dictionary synonyms exist, and the semantically disabled relationship corresponding to the pre-set business scenario is hit, the hit dictionary synonym is deleted; wherein, for each business scenario, the corresponding semantically disabled relationship is set according to the relevant industry dictionary, operation manual, product brochure, and expert experience; the hit here requires the specified dictionary and dictionary synonyms, and hits the two words in a single semantically disabled relationship; If the business scenario belongs to a preset scenario, a third model prompt word corresponding to the dictionary synonym is generated according to a preset synonym screening template; the synonym screening template includes: dictionary enumeration value constraints, belonging field constraints, business scenario constraints, and output format constraints; Based on the third large model prompt word, the large model is called to screen the dictionary synonyms.

7. The conversational analysis retrieval method according to claim 1, characterized in that: Perform analysis and retrieval based on the updated knowledge base, including: Get the search content entered by the user; Perform word segmentation processing on the search content to obtain a number of entity words; According to the entity word, entity information corresponding to relevant dictionaries and relevant fields is recalled in the knowledge base; Determining associated knowledge having an associated relationship in a knowledge base according to the entity information; The entity information and the associated knowledge are converted into a structured language, and the structured language is used for analysis and retrieval.

8. The conversational analysis retrieval method according to claim 7, characterized in that: According to the entity information and the associated knowledge, converting into a structured language specifically includes: Determining a dimension attribute of the entity information; According to the dimension attribute, determining whether the entity information corresponds to a specified field or a specified dictionary; If it corresponds to the specified field, determining that the entity information has a first weight; If it corresponds to the specified dictionary, determining whether the entity information belongs to the original content; If it belongs to the original content, determining the entity information as the second weight; If it does not belong to the original content, determining whether the entity information belongs to the synonym content temporarily obtained this time; If it is a temporarily obtained synonym content, the entity information is determined to be the third weight; If it does not belong to the temporarily obtained synonym content, determining the update time of the entity information in the knowledge base, and determining the fourth weight corresponding to the entity information according to the update time; According to the weight corresponding to the entity information, the entity information and the associated knowledge are converted into a structured language; Among them, the weight values ​​are from high to low: first weight, second weight, third weight, fourth weight.

9. A conversational analysis and retrieval device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the conversational analysis retrieval method as described in any one of claims 1 to 8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured as: the conversational analysis retrieval method described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Medical data filling method and device

    CN114065936A

  • Multi-stage mixed retrieval method and system

    CN119322834A