Conversational analysis retrieval method and device and medium
By generating big model prompt words and calling big model to generate synonyms, the matching problem of dialogue analysis system in relational data retrieval is solved, automatic synonyms addition and knowledge acquisition are realized, and the system's intelligence level and user experience are improved.
Patent Information
- Application Number
- CN202510429242.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing dialogue analysis system based on large models is difficult to accurately match the user's colloquialization problems in relational data retrieval, and the user has not clearly pointed out the required business knowledge, which makes the knowledge unable to be known by the large model, and manual configuration of knowledge is inefficient.
By obtaining data structure information, generating big model prompt words, calling big model to generate fields and dictionary synonyms, performing knowledge retrieval and association, and updating the knowledge base to support analysis retrieval.
It reduces manual intervention, reduces labor cost investment, improves the accuracy of synonyms and the efficiency of obtaining business knowledge, and promotes the intelligent development of dialogue analysis systems.
Smart Images

Figure CN119938890A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and specifically to conversational analysis retrieval methods, devices and media. Background Art
[0002] With the development of technology, large models have gradually come into view.
[0003] In traditional solutions, in conversational analysis systems based on large models, the following key technical issues need to be addressed for relational data retrieval: 1. The original names of fields and dimension dictionaries in existing datasets are mostly configured Chinese names, while user questions are often presented in a colloquial form. This makes it difficult to accurately match the correct fields and the values of enumeration items in the dictionary when performing retrieval-augmented generation (RAG) retrieval, and it is therefore difficult to put the necessary knowledge into the prompt words and transmit them to the large model.
[0004] 2. The user's question may not clearly indicate the need for certain business knowledge, which makes this knowledge unknown to the big model. In this case, although the knowledge can be configured in advance on the data set fields or certain dictionary items, it can be automatically brought out after the recall. However, the input of these field-related business knowledge is labor-intensive, not only inefficient, but also prone to omissions, making it difficult to meet large-scale data processing and real-time requirements. Summary of the invention
[0005] In order to solve the above problems, this application proposes a conversational analysis retrieval method, including: For the preset data set, obtain the corresponding data structure information; For the designated field and / or designated dictionary in the data structure information, as the original content, and through a preset synonym generation template, a first large model prompt word corresponding to the original content is generated; Based on the first large model prompt word, calling the large model, generating corresponding field synonyms and / or dictionary synonyms as synonym content; Searching the synonymous content in the knowledge base, recalling the corresponding business knowledge, and establishing an association relationship between the business knowledge and the original content; Based on the association relationship, the knowledge base is updated, and analysis and retrieval are performed based on the updated knowledge base.
[0006] In one example, the first large model prompt word corresponding to the original content is generated by using a preset synonym generation template, specifically including: For the specified field, determine the corresponding field synonym generation template; the field synonym generation template includes: field name constraints, data set business purpose constraints, field type constraints, and output format constraints; For the specified dictionary, determine the corresponding dictionary synonym generation template; the dictionary synonym generation template includes: dictionary enumeration value constraints, belonging field constraints, description length constraints, superior-subordinate relationship constraints, and output format constraints; According to the field synonym generation template and the dictionary synonym generation template, corresponding first large model prompt words are generated for the designated field and the designated dictionary respectively.
[0007] In one example, based on the first large model prompt word, calling the large model, generating corresponding field synonyms and / or dictionary synonyms as synonym content, the method further includes: Determine, according to the data structure information, the superior-subordinate relationship corresponding to the designated dictionary; Perform a duplicate check based on the dictionary synonyms to determine whether the same dictionary synonyms exist in different designated dictionaries; If it exists, the dictionary synonym is used as the conflicting synonym; The conflicting synonyms are deduplicated according to the superior-subordinate relationship of the designated dictionaries corresponding to the conflicting synonyms.
[0008] In one example, according to the superior-subordinate relationship of the designated dictionaries corresponding to the conflicting synonyms, deduplication of the conflicting synonyms specifically includes: Determining a plurality of designated dictionaries corresponding to the conflicting synonyms; If there is a superior-subordinate relationship among the multiple designated dictionaries, the dictionary synonyms corresponding to the top-level designated dictionary are retained; If the multiple designated dictionaries are in a peer relationship, then according to the order of the designated dictionaries, the dictionary synonyms corresponding to the frontmost designated dictionary are retained.
[0009] In one example, after determining a plurality of designated dictionaries corresponding to the conflicting synonyms, the method further includes: Determining whether there is a preset core term in the plurality of designated dictionaries or the conflicting synonyms; If a core term exists in the multiple specified dictionaries, the superior-subordinate relationship is ignored, and the dictionary synonyms corresponding to the core term are retained; If the conflicting synonym is a core term, a second largest model prompt word corresponding to the core term is generated according to a preset proximity determination template; the proximity determination template includes: dictionary enumeration value constraints, belonging field constraints, core term constraints, and output format constraints; Based on the second large model prompt word, the large model is called to determine the similarity between the core term and each designated dictionary, and the designated dictionary corresponding to the core term to be retained is determined according to the similarity and the hierarchical relationship corresponding to the multiple designated dictionaries.
[0010] In one example, based on the first large model prompt word, calling the large model, generating corresponding field synonyms and / or dictionary synonyms as synonym content, the method further includes: Determine the designated dictionary and the corresponding dictionary synonyms, and determine the business scenario according to the designated dictionary; If the designated dictionary and the dictionary synonym exist, and the semantically disabled relationship corresponding to the preset business scenario is hit, the hit dictionary synonym is deleted; If the business scenario belongs to a preset scenario, a third model prompt word corresponding to the dictionary synonym is generated according to a preset synonym screening template; the synonym screening template includes: dictionary enumeration value constraints, belonging field constraints, business scenario constraints, and output format constraints; Based on the third large model prompt word, the large model is called to screen the dictionary synonyms.
[0011] In one example, performing analysis and retrieval based on the updated knowledge base specifically includes: Get the search content entered by the user; Perform word segmentation processing on the search content to obtain a number of entity words; According to the entity word, entity information corresponding to relevant dictionaries and relevant fields is recalled in the knowledge base; Determining associated knowledge having an associated relationship in a knowledge base according to the entity information; The entity information and the associated knowledge are converted into a structured language, and the structured language is used for analysis and retrieval.
[0012] In one example, converting into a structured language according to the entity information and the associated knowledge specifically includes: Determining a dimension attribute of the entity information; According to the dimension attribute, determining whether the entity information corresponds to a specified field or a specified dictionary; If it corresponds to the specified field, determining that the entity information has a first weight; If it corresponds to the specified dictionary, determining whether the entity information belongs to the original content; If it belongs to the original content, determining the entity information as the second weight; If it does not belong to the original content, determining whether the entity information belongs to the synonym content temporarily obtained this time; If it is a temporarily obtained synonym content, the entity information is determined to be the third weight; If it does not belong to the temporarily obtained synonym content, determining the update time of the entity information in the knowledge base, and determining the fourth weight corresponding to the entity information according to the update time; According to the weight corresponding to the entity information, the entity information and the associated knowledge are converted into a structured language; Among them, the weight values are from high to low: first weight, second weight, third weight, fourth weight.
[0013] On the other hand, the present application also proposes a conversational analysis and retrieval device, comprising: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the conversational analysis retrieval method described in any of the above examples.
[0014] On the other hand, the present application also proposes a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to be: the conversational analysis retrieval method described in any of the above examples.
[0015] The conversational analysis retrieval method proposed in this application can bring the following beneficial effects: 1. Compared with the traditional method of manually adding synonyms and knowledge, it does not require a lot of manpower and time costs. The large model can be used to add synonyms and acquire knowledge, which reduces manual intervention and labor cost investment, thereby increasing work efficiency while reducing costs.
[0016] 2. Compared with traditional manual addition, the use of big model technology can generate more accurate synonyms that are more in line with business scenarios, thereby increasing the accuracy of the knowledge finally returned.
[0017] 3. By deeply integrating big model technology into the relational data retrieval process, the conversational analysis system has been promoted to develop in an intelligent direction. The function of automatically generating synonyms allows the system to better simulate the way human language is understood and interact with users more naturally, which improves the intelligence level and user experience of the entire system and lays the foundation for the intelligent upgrade of related fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A flowchart of a conversational analysis and retrieval method in an embodiment of the present application; Figure 2 A schematic diagram of a conversational analysis retrieval method in one scenario in an embodiment of the present application; Figure 3 It is a schematic diagram of the conversational analysis retrieval device in an embodiment of the present application. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0020] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.
[0021] To solve the above-mentioned problems, we can build an algorithm around probability models and computational logic to generate synonyms. For example, we can set up a statistical algorithm to determine synonyms by analyzing the co-occurrence frequency of words in a large-scale text corpus. The core of this algorithm is to determine which words have similar semantics based on the probability relationship of the words appearing in the text. There are also algorithms based on deep learning, which use neural network models to learn the vector representation of words and identify synonyms by calculating the similarity between vectors.
[0022] However, in data analysis scenarios, the situation may become more complicated. Although they are all synonyms, there may be differences between the dataset field context, dimension table dictionary context, and business knowledge context. Taking the dataset field context as an example, each field has its specific data meaning and purpose. For example, in financial data analysis, the "sales" field represents the amount of revenue obtained by the company through the sale of products or services. The synonyms in this context must closely revolve around the business concept, and the personalized difference knowledge of these concepts in a specific company must also be delivered to the big model. The dimension table dictionary context focuses on the definition and interpretation of data dimensions. For example, in a dimension table dictionary containing a regional dimension, "province" may correspond to the name of a specific provincial administrative region, and the generation of its synonyms must be consistent with the context of this dimension definition. The business knowledge context is more macro, covering the professional knowledge, industry rules, and business processes of the entire business field. For example, in the medical data analysis scenario, the business knowledge context contains information such as medical terms, disease diagnosis standards, and treatment processes. The generation of synonyms must conform to these professional knowledge backgrounds.
[0023] However, traditional synonym generation algorithms often have difficulty making full use of this additional information. Since most of these algorithms are trained based on general natural language texts and are not optimized for the special contexts of data analysis scenarios, when processing synonyms related to data set fields, dimension table dictionaries, and business knowledge, the generated synonyms may not be accurate or may not meet actual business needs. For example, in general texts, "spend" and "expenditure" may be considered synonyms, but in the context of financial data analysis, "spend" may focus more on the concept of daily consumption, while "expenditure" is more financially professional and suitable for formal financial statement records. Existing general algorithms may find it difficult to accurately distinguish this contextual difference.
[0024] Based on this, the solution for private domain knowledge can establish a private domain knowledge base, and recall relevant knowledge from the knowledge base at runtime, and then perform precise sorting or reasoning to eliminate interference items. For example, the user's question mentions an entity, how is a certain indicator this year. During the recall, only the indicator information of the indicator can be recalled according to the question, and the knowledge related to the indicator needs to be recalled again in the knowledge base. Since the knowledge base is retrieved based on the similarity algorithm, this process will increase the probability of making mistakes. Obviously, the second recall can solve the problem by storing the indicator ID and related knowledge in a structured manner. However, it is necessary to associate the indicator ID with the relevant knowledge in advance, which usually relies on manual work, which is time-consuming and labor-intensive.
[0025] Based on this, Figure 1 As shown, the embodiment of the present application provides a conversational analysis and retrieval method to solve the problems existing in the above-mentioned traditional solutions.
[0026] like Figure 1and Figure 2 As shown, the method includes: S101: For a preset data set, obtain corresponding data structure information.
[0027] The conversational analysis and retrieval system (hereinafter referred to as the analysis system) means that the user asks questions to the analysis system, and the analysis system retrieves the corresponding knowledge based on the questions asked by the user, and then feeds back the corresponding answers.
[0028] The analysis system generally writes one or more complete query SQL statements for the subject to be analyzed to form a queryable data view, which can be called a data set.
[0029] The analysis system is usually associated with a business intelligence system (BI system) or is a system module in the BI system. The BI system can automatically identify the data field type, making it easier for the analysis system to perform corresponding analysis and retrieval.
[0030] And implementers usually manually mark the name of each field, manually give the dataset a name that can represent the meaning of the dataset, add a dataset description to the dataset, and make notes to describe the data contained.
[0031] Based on this, the data set is traversed by topic to obtain the corresponding data structure information. The data structure information may include: a brief description of the topic, the name of the data set, the description of the data set, the name of each field in the data set, the field type information, and the dimension dictionary information associated with the field. According to the dimension dictionary information, the text description of each enumeration item in each dimension dictionary, as well as the level and superior-subordinate relationship are read. Among them, the superior-subordinate relationship refers to the semantic relationship of inclusion and being included. For example, it includes: the geographical inclusion of Province A and City B within Province A, the semantic inclusion of total sales and sales of a certain product, etc. Of course, the superior-subordinate relationship here mainly refers to the superior-subordinate relationship within the specified dictionary. If a specified dictionary has a superior relationship, but the superior relationship does not belong to the specified dictionary, it does not need to be considered.
[0032] S102: For a designated field and / or designated dictionary in the data structure information, as original content, and using a preset synonym generation template, generate a first large model prompt word corresponding to the original content.
[0033] The specified field and the specified dictionary refer to the part or all of the original fields and dictionaries in the data table selected according to the requirements or settings. The dictionary can also be called the dimension dictionary or dimension table dictionary, which refers to the specific value of a piece of data under the field.
[0034] Specifically, for a specified field, determine its corresponding field synonym generation template (which can also be called a field synonym Prompt template); the field synonym generation template includes: field name constraint, dataset business purpose constraint, field type constraint, and output format constraint.
[0035] Generally speaking, the field synonym generation templates for each field are the same, but different templates can also be set based on requirements.
[0036] For example, the field synonym generation template can be: <!Field Name!> is a column in the data table <!Dataset Name!>, the business purpose of the table is <!Dataset Description!>, the field type is <!Field Type!>, please generate possible synonyms for the field and return them in the form of a JSON array.
[0037] Among them, <!String!> means the system variable corresponding to this "string", which can be replaced with the content in the data structure information. At this time, the variable corresponding to the field name corresponds to the field name constraint, the variable corresponding to the dataset description corresponds to the dataset business purpose constraint, the variable corresponding to the field type corresponds to the field type constraint, and returning in the form of a JSON array corresponds to the output format constraint.
[0038] Similarly, for a specified dictionary, determine its corresponding dictionary synonym generation template; the dictionary synonym generation template includes: dictionary enumeration value constraint, belonging field constraint, description length constraint, superior-subordinate relationship constraint, and output format constraint.
[0039] For example, the dictionary synonym generation template can be: <!Dictionary Enumeration Value!> is one of the values in the column <!Field Name!> in the data table <!Dataset Name!>, the field length is <!Field Length!>, its superior is <!Superior Name!>, and it has <!Number of Subordinates!> subordinates. Please generate possible synonyms and return them in the form of a JSON array.
[0040] Among them, the dictionary enumeration value refers to the specific value corresponding to each specified dictionary under each field. The variable corresponding to the dictionary enumeration value corresponds to the dictionary enumeration value constraint, the variable corresponding to the field name corresponds to the belonging field constraint, the variable corresponding to the field length (here it refers to the length of this dictionary) corresponds to the description length constraint, the variables corresponding to the superior name and the number of subordinates correspond to the superior-subordinate relationship constraint, and returning in the form of a JSON array corresponds to the output format constraint.
[0041] Based on this, according to the field synonym generation template and the dictionary synonym generation template, generate the corresponding first large model prompt words for the specified field and the specified dictionary respectively.
[0042] For example, the first model prompt word corresponding to the specified field can be: Product name is a column in the data table e-commerce sales status table. The business purpose of the table is to describe the sales and sales volume of each product on the e-commerce platform within a certain period of time. The field type is varchar. Please generate possible synonyms for the field and return them in the form of a JSON array.
[0043] The first model prompt word corresponding to the specified dictionary can be: Tablet computer is one of the values in a column of product names in the data table e-commerce sales status table, the field length is varchar(10), its superior is computer, and it has 10 subordinates, please generate possible synonyms. And return it in JSON array form.
[0044] S103: Based on the first large model prompt word, call the large model to generate corresponding field synonyms and / or dictionary synonyms as synonym content.
[0045] At this time, for each first large model prompt word, the large model API is called to generate and return corresponding synonym content through the large model. For the same original content, it can include one or more synonym content. In some scenarios, it can also be set to no corresponding synonym content found.
[0046] The obtained field synonyms can be recorded in the field information corresponding to the specified field, and the obtained dictionary synonyms can also be recorded in the corresponding dictionary name structure.
[0047] Furthermore, after obtaining the synonym content, since the specified fields and the specified dictionary are all from the same data table, it is very likely that the same synonym content exists. For example, there are two specified dictionaries, namely "tablet computer" and "desktop computer", which may both contain the same synonym "computer".
[0048] At this time, if the synonym content is used directly, it is easy to cause content duplication, which not only wastes resources but also easily causes subsequent content conflicts. Therefore, deduplication processing can be performed after the synonym content is obtained.
[0049] According to the data structure information, the superior-subordinate relationship corresponding to the specified dictionary is determined. As mentioned above, the superior-subordinate relationship corresponding to each specified dictionary has been determined in the data structure information. In the superior-subordinate relationship, multiple levels can be pre-set, and the level of each specified dictionary is set to obtain the corresponding superior-subordinate relationship.
[0050] For example, there are three levels. The first level is the highest level, which includes two designated dictionaries, namely "Area A" and "Computer", and the second level is the middle level, which includes "Area B", "Area C", "Tablet Computer", "Desktop Computer", etc., among which "Area B" and "Area C" both belong to "Area A". At this time, "Area B" and "Area C" are in a peer relationship, and they are in a subordinate relationship with "Area A". Similarly, "Tablet Computer" and "Desktop Computer" are also in a peer relationship, and they are in a subordinate relationship with "Computer". Of course, in order to ensure the clarity of the superior-subordinate relationship, the corresponding connection relationship can be set to clarify which designated dictionaries have a superior-subordinate relationship between different levels.
[0051] Check for duplicates based on dictionary synonyms to determine whether the same dictionary synonyms exist in different specified dictionaries.
[0052] If it does not exist, no deduplication processing is required.
[0053] If it exists, the dictionary synonym is used as a conflicting synonym. At this time, the conflicting synonyms are deduplicated according to the upper and lower relationships of the designated dictionaries corresponding to the conflicting synonyms.
[0054] Specifically, when removing duplicates, multiple designated dictionaries corresponding to conflicting synonyms are determined.
[0055] If there is a hierarchical relationship among multiple designated dictionaries, the dictionary synonyms corresponding to the top designated dictionary are retained. That is, the conflicting synonyms are used as the dictionary synonyms corresponding to the top designated dictionary, and the conflicting synonyms corresponding to other designated dictionaries are deleted.
[0056] If multiple specified dictionaries are of the same level, the dictionary synonyms corresponding to the first specified dictionary are retained according to the order of the specified dictionaries.
[0057] Generally speaking, only the specified dictionaries with a hierarchical relationship (including superior, subordinate, and peer) can have the same synonyms. Of course, in rare cases, there may be conflicting synonyms between two dictionaries that do not have a hierarchical relationship. In this case, the hierarchical relationship between the two dictionaries can be determined based on the level of the dictionary.
[0058] For the specified dictionaries with a superior or subordinate relationship, the dictionary synonyms corresponding to the top-level specified dictionary are retained. The conflicting synonyms can be stored later to prevent the subordinate from repeatedly using the superior name, ensure the integrity of the category data, and prevent unnecessary data pollution. For example, assuming that in the specified dictionary, "Product D" is the superior relationship of "Product E", and the conflicting synonym corresponding to the two is "Product F". At this time, there may be three situations for Product F, namely: the first, Product F is a higher level of Product D, the second, Product F and Product D belong to the same level, and the third, Product F and Product E belong to the same level or a lower level.
[0059] At this time, for the first case, whether the conflicting synonym "Product F" is retained as a synonym of "Product D" or as a synonym of "Product E", when the user searches for Product D or Product E, it will cause products that do not belong to themselves to appear, thereby causing data pollution. However, at this time, since Product D is at a higher level and covers a larger range, retaining the conflicting synonym "Product F" as a synonym of "Product D" will cause relatively less data pollution.
[0060] Of course, if the various levels in the superior-subordinate relationship are reasonably divided in advance, the first situation will usually not occur.
[0061] In the second case, when the conflicting synonym is retained as a synonym of "product D", since the two are at the same level, there will be no pollution when the user searches for product D. However, when the conflicting synonym is retained as a synonym of "product E", since the level of product F is higher than that of product E, when the user searches for product E, data pollution will still occur, and products belonging to product F but not product E will appear.
[0062] As for the third case, there will be no data pollution in either case.
[0063] Based on this, retaining the dictionary synonyms corresponding to the top-level designated dictionary can increase the user's subsequent search range while reducing data pollution. For designated dictionaries at the same level, retaining dictionary synonyms for any of them has basically similar effects, so they can be randomly selected or retained in the order of appearance.
[0064] S104: Searching the synonymous content in the knowledge base, recalling the corresponding business knowledge, and establishing an association relationship between the business knowledge and the original content.
[0065] Through the business rules stored in the knowledge base, the business knowledge corresponding to the synonym content can be retrieved to obtain the required business knowledge. For each synonym content, a separate search is performed, and then the retrieved business knowledge is associated with the original content corresponding to the synonym content.
[0066] The association relationship here usually refers to the association relationship stored in the cache. Through this association relationship, when the user searches through the original content, he can retrieve the business knowledge associated with the corresponding synonym content.
[0067] S105: Based on the association relationship, the knowledge base is updated, and analysis and retrieval are performed based on the updated knowledge base.
[0068] In order to ensure that the business knowledge corresponding to the synonym content can still be retrieved through the association relationship in the subsequent process, the knowledge base can be updated according to the association relationship and the association relationship can be written into the knowledge base, so that users can also retrieve relevant business knowledge in the subsequent search process.
[0069] The update may be performed in a regular manner, for example, the previously obtained association relationship is temporarily stored in a cache, and after a certain period of time, the association relationship in the cache is updated to the knowledge base, and the corresponding update timestamp is recorded.
[0070] 1. Compared with the traditional method of manually adding synonyms and knowledge, it does not require a lot of manpower and time costs. The large model can be used to add synonyms and acquire knowledge, which reduces manual intervention and labor cost investment, thereby increasing work efficiency while reducing costs.
[0071] 2. Compared with traditional manual addition, the use of big model technology can generate more accurate synonyms that are more in line with business scenarios, thereby increasing the accuracy of the knowledge finally returned.
[0072] 3. By deeply integrating big model technology into the relational data retrieval process, the conversational analysis system has been promoted to develop in an intelligent direction. The function of automatically generating synonyms allows the system to better simulate the way human language is understood and interact with users more naturally, which improves the intelligence level and user experience of the entire system and lays the foundation for the intelligent upgrade of related fields.
[0073] In one embodiment, when performing deduplication of conflicting synonyms, there may be some special cases, so that the final effect is not satisfactory if only the deduplication method described above is used.
[0074] Specifically, for some special commodities or terms in special industries (such as the medical industry), there are some words that are core words. Although their levels may not be high, they are relatively more important.
[0075] At this time, determine whether there is a preset core term in multiple specified dictionaries or conflict synonyms. The core term can be obtained through relevant industry dictionaries, operation manuals, product brochures, expert experience, etc.
[0076] If there is a core term in multiple specified dictionaries, ignore the hierarchical relationship and retain the dictionary synonyms corresponding to the core term. Even if the level of the core term is low, the hierarchical relationship is no longer considered, the dictionary synonyms corresponding to the core term are retained, and the dictionary synonyms corresponding to other specified dictionaries are deleted, so as to highlight the importance of the core term.
[0077] If the conflict synonym is a core term, considering the particularity of the core term at this time, only relying on the hierarchical relationship may cause confusion in the finally retained synonym relationship.
[0078] Therefore, according to the preset approximation determination template, generate the second large model prompt word corresponding to the core term; the approximation determination template includes: dictionary enumeration value constraint, belonging field constraint, core term constraint, output format constraint.
[0079] For example, the approximation determination template can be: <!Dictionary enumeration value!> and <!Dictionary enumeration value!> are both one of the values in a column <!Field name!> in the data table <!Dataset name!>. Please judge the similarity between it and <!Core term enumeration value!> and output the similarity in percentage form. And return it in the form of a JSON array.
[0080] Among them, the approximation determination template can be set to multiple, and the number of <!Dictionary enumeration value!> in each template is different, which is used to correspond to the number of specified dictionaries.
[0081] Based on the second large model prompt word, call the large model to determine the similarity between the core term and each specified dictionary.
[0082] According to the similarity and the hierarchical relationship corresponding to multiple specified dictionaries, determine the specified dictionary corresponding to the core term to be retained. For example, set corresponding weights for the similarity and the hierarchical relationship, obtain the corresponding score through weighted summation, and select the specified dictionary with a higher score, and retain the core term as the dictionary synonym. Among them, the score corresponding to the hierarchical relationship can be set to correspond to a score for each difference in level, so as to obtain the score corresponding to the hierarchical relationship according to the level of each specified dictionary, and perform weighted summation with the score corresponding to the similarity to obtain the final score.
[0083] In one embodiment, in addition to the special cases corresponding to the core terms mentioned above, there may be some other special cases that affect the deduplication effect.
[0084] Specifically, for the special requirements of some merchants or the special requirements of the industry, in order to ensure the uniqueness of their own products, or for the industry to require the independence of certain terms, it is prohibited that there is an associated relationship between certain words and some common synonyms.
[0085] Based on this, a specified dictionary and its corresponding dictionary synonyms are determined, and the business scenario to which it belongs is determined according to the specified dictionary. The business scenario can be obtained by whether the specified dictionary hits relevant keywords. For example, some common keywords are pre-selected and relationships are established with each business scenario.
[0086] For each business scenario, corresponding semantic disabling relationships can be set according to relevant industry dictionaries, operation manuals, product brochures, expert experience, etc. Some of these disabling relationships prohibit certain words from being synonymous. Setting different business scenarios with different semantic disabling relationships can reduce the complexity of comparison and duplicate checking and increase efficiency in the case of a large number of semantic disabling relationships.
[0087] At this time, if there is a specified dictionary and dictionary synonyms that hit the semantic disabling relationship corresponding to the pre-set business scenario, the hit dictionary synonyms will be deleted. Here, for a hit, both the specified dictionary and the dictionary synonyms need to hit two words in a single semantic disabling relationship. If only one of the specified dictionary and the dictionary synonyms hits one word in a certain semantic disabling relationship, it is not considered a hit of that semantic disabling relationship.
[0088] If the business scenario belongs to a preset scenario (such as industries with strict requirements like the medical industry and the legal industry), then in addition to setting this semantic disabling relationship, third-party model prompt words corresponding to the dictionary synonyms can also be generated according to a pre-set synonym screening template; the synonym screening template includes: dictionary enumeration value constraints, field constraints, business scenario constraints, and output format constraints.
[0089] For example, the synonym screening template can be: <!Dictionary enumeration value!> is one of the values in the column <!Field name!> of the data table <!Dataset name!>, and it has the following synonyms: <!Synonym enumeration value!>, <!Synonym enumeration value!>. The business scenario of this <!Dictionary enumeration value!> is <!Business scenario name!>. Please screen out among the above synonyms and only retain the synonyms that conform to this business scenario. And return it in the form of a JSON array.
[0090] Among them, multiple synonym screening templates can be set, and the number of <!synonym enumeration values!> in each template is different, corresponding to the number of dictionary synonyms.
[0091] Based on the third large model prompt, the large model is called to screen the dictionary synonyms, so as to ensure the rationality of synonyms in some important industries and businesses.
[0092] In one embodiment, when the user performs an analysis and retrieval through the updated knowledge base, the retrieval content input by the user is obtained. The retrieval content is segmented to obtain a number of entity words. The segmentation process can be implemented by rule-based methods such as forward maximum matching and reverse maximum matching, or can be implemented by traditional probability models such as hidden Markov models and conditional random fields, or can be implemented by deep learning models such as BERT and BiLSTM-CRF.
[0093] According to the entity words, the entity information corresponding to the relevant dictionaries and relevant fields is recalled in the knowledge base. The entity information can include the entity information corresponding to the fields and dictionaries.
[0094] According to the entity information, the associated knowledge with an association relationship is determined in the knowledge base. At this time, the associated knowledge can include the original associated knowledge corresponding to the entity words, or can include the business knowledge with an association relationship brought by the synonyms after the knowledge base is updated.
[0095] According to the entity information and the associated knowledge, it is converted into a structured language, and the analysis and retrieval are performed through the structured language. At this time, it can be input into the NL2SQL agent to improve the accuracy. Among them, the NL2SQL agent is an intelligent system based on artificial intelligence technology, and its core function is to automatically convert the natural language instructions input by the user into executable SQL statements, so as to achieve efficient query and management of the database.
[0096] Furthermore, when converting to a structured language, due to the business knowledge in the knowledge base, it may include the business knowledge directly obtained from the original content and the business knowledge obtained from the synonym content in different time periods, resulting in a large amount of business knowledge data. If directly used, it may lead to chaos in the business knowledge, and finally the retrieval results returned to the user do not fully meet the user's needs.
[0097] Based on this, the dimension attributes of the entity information are determined. Its dimension attributes may include: belonging to a field or dictionary, whether it belongs to a synonym, whether it belongs to a synonym temporarily obtained this time, the update time of the synonym in the knowledge base, etc. For example, for a certain entity information, if it is a newly added designated dictionary for the user's current search, its dimension attributes are: [dictionary; belongs to; belongs to; empty]. For the convenience of description, the dimension attribute can be described by the corresponding code. For example, set the "belongs to" code to 1, set the "does not belong to" code to 0, set the field code to 1, set the dictionary code to 0, etc.
[0098] According to the dimension attributes, determine whether the entity information corresponds to the specified field or the specified dictionary. If it corresponds to the specified field, the entity information is determined to have the first weight. Compared with dictionaries, fields contain more general content and can highlight the essential purpose of the questions asked by users. Therefore, they are judged first, and when they are fields, they can be considered to be set to the highest first weight to indicate the core purpose of the user.
[0099] If it corresponds to the specified dictionary, continue to judge whether the entity information belongs to the original content. If it belongs to the original content, the entity information is determined to be the second weight. The original content is more in line with the user's initial idea than the business knowledge obtained from the synonym content, so if it belongs to the original content of the specified dictionary, it is set to a second weight lower than the first weight.
[0100] If it does not belong to the original content, continue to judge and determine whether the entity information belongs to the temporarily obtained synonym content. If it belongs to the temporarily obtained synonym content, the entity information is determined to be the third weight. Compared with the content that has been updated in the knowledge base before, the temporarily obtained synonym content is more in line with the current scenario requirements and can better ensure the timeliness of the synonyms. Therefore, if it belongs to the temporarily obtained synonym content of the specified dictionary, a third weight lower than the second weight is set for it.
[0101] If it does not belong to the synonymous content obtained temporarily, the update time of the entity information in the knowledge base is determined, and the fourth weight corresponding to the entity information is determined according to the update time. Similarly, if it belongs to the synonymous content that has been updated in the knowledge base, the fourth weight lower than the third weight can be set according to the update time. Among them, the closer the update time is to the current time, the higher the fourth weight is, so as to ensure the timeliness of the entity information.
[0102] At this time, the entity information and associated knowledge are converted into structured language according to the weight corresponding to the entity information, so as to ensure that the data finally returned meets the user's needs.
[0103] like Figure 3As shown, the embodiment of the present application also provides a conversational analysis and retrieval device, including: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the conversational analysis retrieval method described in any of the above examples.
[0104] An embodiment of the present application also provides a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to be: the conversational analysis retrieval method described in any of the above examples.
[0105] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0106] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0107] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A conversational analysis retrieval method, characterized in that: include: For the preset data set, obtain the corresponding data structure information; For the designated field and / or designated dictionary in the data structure information, as the original content, and through a preset synonym generation template, a first large model prompt word corresponding to the original content is generated; Based on the first large model prompt word, calling the large model, generating corresponding field synonyms and / or dictionary synonyms as synonym content; Searching the synonymous content in the knowledge base, recalling the corresponding business knowledge, and establishing an association relationship between the business knowledge and the original content; Based on the association relationship, the knowledge base is updated, and analysis and retrieval are performed based on the updated knowledge base.
2. The conversational analysis retrieval method according to claim 1, characterized in that: Generate the first model prompt word corresponding to the original content by using a preset synonym generation template, specifically including: For the specified field, determine the corresponding field synonym generation template; the field synonym generation template includes: field name constraints, data set business purpose constraints, field type constraints, and output format constraints; For the specified dictionary, determine the corresponding dictionary synonym generation template; the dictionary synonym generation template includes: dictionary enumeration value constraints, belonging field constraints, description length constraints, superior-subordinate relationship constraints, and output format constraints; According to the field synonym generation template and the dictionary synonym generation template, corresponding first large model prompt words are generated for the designated field and the designated dictionary respectively.
3. The conversational analysis retrieval method according to claim 1, characterized in that: Based on the first large model prompt word, calling the large model to generate corresponding field synonyms and / or dictionary synonyms as synonym content, the method further includes: Determine, according to the data structure information, the superior-subordinate relationship corresponding to the designated dictionary; Perform a duplicate check based on the dictionary synonyms to determine whether the same dictionary synonyms exist in different designated dictionaries; If it exists, the dictionary synonym is used as the conflicting synonym; The conflicting synonyms are deduplicated according to the superior-subordinate relationship of the designated dictionaries corresponding to the conflicting synonyms.
4. The conversational analysis retrieval method according to claim 3, characterized in that: Deduplication of the conflicting synonyms is performed according to the superior-subordinate relationship of the designated dictionaries corresponding to the conflicting synonyms, specifically including: Determining a plurality of designated dictionaries corresponding to the conflicting synonyms; If there is a superior-subordinate relationship among the multiple designated dictionaries, the dictionary synonyms corresponding to the top-level designated dictionary are retained; If the multiple designated dictionaries are in a peer relationship, then according to the order of the designated dictionaries, the dictionary synonyms corresponding to the frontmost designated dictionary are retained.
5. The conversational analysis retrieval method according to claim 4, characterized in that: After determining the multiple designated dictionaries corresponding to the conflicting synonyms, the method further includes: Determining whether there is a preset core term in the plurality of designated dictionaries or the conflicting synonyms; If a core term exists in the multiple specified dictionaries, the superior-subordinate relationship is ignored, and the dictionary synonyms corresponding to the core term are retained; If the conflicting synonym is a core term, a second largest model prompt word corresponding to the core term is generated according to a preset proximity determination template; the proximity determination template includes: dictionary enumeration value constraints, belonging field constraints, core term constraints, and output format constraints; Based on the second large model prompt word, the large model is called to determine the similarity between the core term and each designated dictionary, and the designated dictionary corresponding to the core term to be retained is determined according to the similarity and the hierarchical relationship corresponding to the multiple designated dictionaries.
6. The conversational analysis retrieval method according to claim 1, characterized in that: Based on the first large model prompt word, calling the large model to generate corresponding field synonyms and / or dictionary synonyms as synonym content, the method further includes: Determine the designated dictionary and the corresponding dictionary synonyms, and determine the business scenario according to the designated dictionary; If the designated dictionary and the dictionary synonym exist, and the semantically disabled relationship corresponding to the preset business scenario is hit, the hit dictionary synonym is deleted; If the business scenario belongs to a preset scenario, a third model prompt word corresponding to the dictionary synonym is generated according to a preset synonym screening template; the synonym screening template includes: dictionary enumeration value constraints, belonging field constraints, business scenario constraints, and output format constraints; Based on the third large model prompt word, the large model is called to screen the dictionary synonyms.
7. The conversational analysis retrieval method according to claim 1, characterized in that: Perform analysis and retrieval based on the updated knowledge base, including: Get the search content entered by the user; Perform word segmentation processing on the search content to obtain a number of entity words; According to the entity word, entity information corresponding to relevant dictionaries and relevant fields is recalled in the knowledge base; Determining associated knowledge having an associated relationship in a knowledge base according to the entity information; The entity information and the associated knowledge are converted into a structured language, and the structured language is used for analysis and retrieval.
8. The conversational analysis retrieval method according to claim 7, characterized in that: According to the entity information and the associated knowledge, converting into a structured language specifically includes: Determining a dimension attribute of the entity information; According to the dimension attribute, determining whether the entity information corresponds to a specified field or a specified dictionary; If it corresponds to the specified field, determining that the entity information has a first weight; If it corresponds to the specified dictionary, determining whether the entity information belongs to the original content; If it belongs to the original content, determining the entity information as the second weight; If it does not belong to the original content, determining whether the entity information belongs to the synonym content temporarily obtained this time; If it is a temporarily obtained synonym content, the entity information is determined to be the third weight; If it does not belong to the temporarily obtained synonym content, determining the update time of the entity information in the knowledge base, and determining the fourth weight corresponding to the entity information according to the update time; According to the weight corresponding to the entity information, the entity information and the associated knowledge are converted into a structured language; Among them, the weight values are from high to low: first weight, second weight, third weight, fourth weight.
9. A conversational analysis and retrieval device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the conversational analysis retrieval method as described in any one of claims 1 to 8.
10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured as: the conversational analysis retrieval method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Method for automatically clustering synonym search results according to lexical meanings
CN103049524A
Medical data filling method and device
CN114065936A
Knowledge graph question and answer method and device, computer equipment and storage medium
CN116303923A
Method and device for retrieving power business database table
CN118673041A
Model cue word automatic optimization method and device, equipment and storage medium
CN119226476A