An enterprise data retrieval method, system and medium based on large models

By configuring database priority and vectorized processing data in enterprise data retrieval, and combining large models to perform user demand analysis and data analysis, the problem of insufficient database priority and user demand guidance in the existing technology is solved, and efficient and accurate enterprise data retrieval and result generation are achieved.

CN120067141BActive Publication Date: 2025-06-27SHANDONG LANGCHAO SMART CULTURAL TOURISM IND DEV CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510550147.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-06-27
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The prior art fails to effectively consider the priority differences between different databases in enterprise data retrieval, and lacks an active guidance mechanism for user needs, resulting in unreasonable weight allocation of search results and semantic ambiguity affecting accuracy.

Method used

By creating an enterprise database and configuring the corresponding relationship between the database, search constraint prompt words and preset search priority, the data is vectorized, the user needs data is obtained to extract search constraint prompt words and determine the database priority, and data analysis is performed in combination with the big model to guide data generation, and finally the output data is generated through the big model processing.

Benefits of technology

Dynamic priority configuration is realized, retrieval efficiency and accuracy are improved, user needs analysis capabilities are enhanced, cross-modal data is supported, and the quality and adaptability of result generation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067141B_ABST
    Figure CN120067141B_ABST
Patent Text Reader

Abstract

The present application discloses a method, system and medium for enterprise data retrieval based on a large model, mainly related to the technical field of enterprise data retrieval, and is used to solve the problems existing in the existing solutions that the priority differences of different databases are not considered and there is a lack of an active guidance mechanism for user requirements. It includes: performing vectorization processing on the data in the enterprise database to convert the text-form data into a numerical vector representation; obtaining user requirement data through the user interaction interface, and then extracting retrieval constraint prompt words, and then determining the preset retrieval priorities corresponding to each enterprise database; calculating the vector representation with the highest similarity to the retrieval constraint prompt words with the adjustment weights corresponding to the preset retrieval priorities; performing data analysis on the retrieval constraint prompt words to determine the corresponding guidance data; inputting the guidance data and the vector representation into the large model to obtain output data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of enterprise data retrieval, and particularly to an enterprise data retrieval method, system, and medium based on a large model. Background Art

[0002] In recent years, with the improvement of enterprise informatization, the scale of enterprise data has shown an explosive growth, and the data form has become increasingly diversified, covering structured data (such as database forms) and unstructured data (such as documents, logs, emails, etc.). How to efficiently and accurately extract effective information from massive enterprise data has become an important challenge for enterprise intelligent management and decision-making. Traditional data retrieval technologies mostly rely on keyword matching or rule engines, such as search engines based on inverted indexes or conditional queries with preset rules. However, such methods have significant limitations: on the one hand, keyword matching is difficult to capture the deep semantics of user needs, especially when dealing with complex queries or polysemous words, it is easy to cause misdetection or missed detection; on the other hand, the rule engine relies on manually predefined logic, with poor flexibility and difficulty in adapting to dynamic business scenarios and personalized needs.

[0003] In the prior art, some solutions attempt to introduce natural language processing (NLP) technology to improve the retrieval semantic understanding ability. For example, the user query is vectorized through a pre-trained language model and matched with the database content for similarity. However, such methods still have the following problems: 1) In the scenario of joint retrieval of multiple enterprise databases, the priority differences of different databases are not considered, resulting in unreasonable weight allocation of retrieval results; 2) There is a lack of an active guidance mechanism for user needs. When the user's needs are expressed vaguely, it is easy to affect the retrieval accuracy due to semantic ambiguity. Summary of the Invention

[0004] This application provides an enterprise data retrieval method, system, and medium based on a large model to solve the problems that the existing solutions do not consider the priority differences of different databases and lack an active guidance mechanism for user needs.

[0005] In the first aspect, this application provides an enterprise data retrieval method based on a large model, and the method includes:

[0006] Create a number of enterprise databases and configure the corresponding relationships among the enterprise databases, retrieval constraint prompt words, and preset retrieval priorities;

[0007] Perform vectorization processing on the data in the enterprise databases to convert the text-form data into a numerical vector representation;

[0008] Through the user interface, obtain user demand data, then extract retrieval constraint prompt words, and then determine the preset retrieval priorities corresponding to each enterprise database;

[0009] Calculate the vector representation with the highest similarity to the retrieval constraint prompt words using the adjustment weights corresponding to the preset retrieval priorities;

[0010] Conduct data analysis on the retrieval constraint prompt words to determine the corresponding guiding data; wherein, the guiding data at least includes: role positioning guiding words, ability positioning guiding words, task objectives, and output requirements;

[0011] Input the guiding data and the vector representation into the large model to obtain output data.

[0012] In one implementation manner of the present application, after creating a number of enterprise databases and configuring the preset retrieval priorities of the enterprise databases, the method further includes:

[0013] Modify the data in the enterprise database through the backend interface, and / or modify the corresponding relationship between the enterprise database, the retrieval constraint prompt words, and the preset retrieval priorities.

[0014] In one implementation manner of the present application, perform vectorization processing on the data in the enterprise database to convert the text-form data into a numerical vector representation, specifically including:

[0015] Obtain the first slice information through the preset slice acquisition interface;

[0016] Based on the first slice information, split the data in the enterprise database to obtain a number of split segments;

[0017] Obtain the segmentation granularity, and use the segmentation granularity to further segment the split segments to obtain a number of knowledge fields;

[0018] Convert the text-form knowledge fields into a numerical vector representation.

[0019] In one implementation manner of the present application, through the user interaction interface, obtain the user demand data, then extract the retrieval constraint prompt words, and then determine the preset retrieval priorities corresponding to each enterprise database, specifically including:

[0020] Obtain the user demand data through the user interaction interface;

[0021] Input the user demand data into the large model, extract the key features in the user demand data, and then determine whether the key features belong to the preset trigger conditions;

[0022] When it belongs to the preset trigger conditions, extract the retrieval constraint prompt words from the user demand data;

[0023] Based on the corresponding relationship between the enterprise database, the retrieval constraint prompt words, and the preset retrieval priorities, determine the preset retrieval priorities corresponding to each enterprise database.

[0024] In one implementation of the present application, the vector representation with the highest similarity to the retrieval constraint prompt is calculated using the adjustment weight corresponding to the preset retrieval priority, which specifically includes:

[0025] The initial similarity between the retrieval constraint prompt and the vector representation is calculated through a language similarity algorithm;

[0026] The final similarity is obtained by multiplying the adjustment weight by the initial similarity, and then the vector representation with the highest similarity is obtained.

[0027] In one implementation of the present application, before analyzing the data of the retrieval constraint prompt to determine the corresponding guiding data, the method further includes:

[0028] Configuring the corresponding relationship between the retrieval constraint prompt and the guiding data.

[0029] In a second aspect, the present application provides an enterprise data retrieval system based on a large model, and the system includes:

[0030] The RAG (Retrieval-augmented Generation) module is used to create a number of enterprise databases, and configure the corresponding relationship between the enterprise databases, retrieval constraint prompts, and preset retrieval priorities; perform vectorization processing on the data in the enterprise databases, and convert the data in text form into numerical vector representations;

[0031] The user interaction module is used to obtain user demand data through the user interaction interface;

[0032] The large model module is used to extract retrieval constraint prompts, and then determine the preset retrieval priorities corresponding to each enterprise database; calculate the vector representation with the highest similarity to the retrieval constraint prompt using the adjustment weight corresponding to the preset retrieval priority; analyze the data of the retrieval constraint prompt to determine the corresponding guiding data; where the guiding data at least includes: role positioning guiding words, ability positioning guiding words, task objectives, and output requirements; input the guiding data and the vector representation into the large model to obtain output data.

[0033] In one implementation of the present application, the RAG module includes a revision unit,

[0034] which is used to modify the data in the enterprise database through the backend interface, and / or modify the corresponding relationship between the enterprise database, retrieval constraint prompt, and preset retrieval priority.

[0035] In one implementation of the present application, the large model module includes a determination unit,

[0036] which is used to extract the key features in the user demand data, and then determine whether the key features belong to the preset trigger conditions;

[0037] When it belongs to the preset trigger condition, extract the retrieval constraint prompt words from the user requirement data;

[0038] Based on the correspondence relationship between the enterprise database, the retrieval constraint prompt words, and the preset retrieval priority, determine the preset retrieval priority corresponding to each enterprise database.

[0039] Thirdly, the present application provides a non-volatile computer storage medium, on which computer instructions are stored, and when the computer instructions are executed, a method for retrieving enterprise data based on a large model as described in any one of the above is implemented.

[0040] It can be seen from the above technical solutions that the present application has the following advantages:

[0041] 1. Dynamic priority configuration improves retrieval efficiency and accuracy:

[0042] Through the preset retrieval priority, it is ensured that high-value databases (such as core business systems) occupy a higher weight in the retrieval, avoiding interference from low-priority databases (such as log backups) on key results. Retrieve high-weight databases first, reduce unnecessary computational overhead, and improve the system response speed. The combination of similarity calculation and priority weight makes the retrieval results more in line with the actual business needs. For example, financial data is preferentially displayed rather than non-critical documents.

[0043] 2. Active guidance mechanism enhances the ability to analyze user requirements:

[0044] Existing solutions lack active guidance for user requirements and are prone to retrieval deviation due to vague requirements. Through role positioning (such as "financial analyst") and ability positioning (such as "data visualization"), the user's identity and skill background are clarified, reducing semantic ambiguity. The introduction of task objectives (such as "generate quarterly reports") and output requirements (such as "Excel format") transforms vague requirements into executable retrieval conditions. Generate differentiated guidance strategies for different user roles (such as sales, R & D) to improve the adaptability of the retrieval experience.

[0045] 3. Cross-modal vectorization processing unifies the data representation space:

[0046] Existing solutions have insufficient capabilities for processing unstructured data and are difficult to unify the semantic space. Vectorization representation breaks through the text form limitation and supports semantic matching of cross-modal data (such as documents, tables, logs). Numerical vectors can be directly input into the large model for reasoning. For example, relevant content in different databases can be associated through vector similarity calculation. Support multi-condition combined queries (such as "find product defect records in customer complaints") to improve retrieval flexibility.

[0047] 4. Collaborative optimization of large models improves the quality of result generation:

[0048] Existing solutions rely on a single retrieval model and are difficult to handle complex tasks. The large model combines the task objectives in the guiding data (such as "risk assessment report") and the vector retrieval results to generate structured and highly readable outputs (such as reports, charts). The large model can correct potential biases in the retrieval results (such as similarity calculation errors) and supplement missing information through context reasoning. It supports generating diverse output forms according to user needs (such as text summaries, visual charts), improving the practicality of the results. Brief Description of the Drawings

[0049] To more clearly illustrate the technical solutions of the present invention, the accompanying drawings required in the description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0050] Figure 1 It is a flowchart of a method for retrieving enterprise data based on a large model provided by an embodiment of the present application.

[0051] Figure 2 It is a schematic diagram of the internal structure of a system for retrieving enterprise data based on a large model provided by an embodiment of the present application. Detailed Embodiments

[0052] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0053] Those skilled in the art should understand that the embodiments described below are only the preferred embodiments of the present disclosure, and do not mean that the present disclosure can only be implemented through these preferred embodiments. The preferred embodiments are only used to explain the technical principles of the present disclosure, rather than to limit the protection scope of the present disclosure. Based on the preferred embodiments provided by the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts should still fall within the protection scope of the present disclosure.

[0054] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0055] This application builds a knowledge retrieval Agent based on the RAG architecture of large models on the open-source Dify platform. Its basic idea is to use RAG technology to build a knowledge database, improve the retrieval accuracy by designing a text vectorization processing method adapted to knowledge retrieval, and at the same time configure roles, tasks and retrieval capabilities for the large model according to the actual needs of the query.

[0056] The technical solutions proposed in the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0057] The embodiment provides an enterprise data retrieval method based on large models, as Figure 1 shown, the method provided in the embodiment of the present application mainly includes the following steps:

[0058] Step 110, create a number of enterprise databases, and configure the corresponding relationships between the enterprise databases, retrieval constraint prompt words and preset retrieval priorities.

[0059] In some embodiments, after creating a number of enterprise databases and configuring the preset retrieval priorities of the enterprise databases, the method further includes:

[0060] Modify the data in the enterprise database through the backend interface, and / or modify the corresponding relationships between the enterprise databases, retrieval constraint prompt words and preset retrieval priorities.

[0061] This step supports users to create and dynamically update enterprise databases as needed. There can be multiple enterprise databases, and the retrieval priorities of different databases can be configured and adjusted by users in the interaction interface or backend interface of the Agent.

[0062] Those skilled in the art can understand that this step enables enterprises to create dedicated databases according to different business modules (such as customer management, supply chain, finance), and control the retrieval weights through preset priorities. For example, in the e-commerce scenario, the "order database" is set with the highest priority to ensure the timeliness of transaction data retrieval; while the "user portrait database" has the second highest priority and is used to assist the recommendation algorithm.

[0063] In a dynamic adjustment scenario, if a temporary promotion activity is carried out, the priority of the "inventory database" can be raised to the top through the backend interface to prioritize the processing of inventory query requests and avoid overselling problems.

[0064] By binding the retrieval constraint prompt words (such as "region = North China", "product category = electronic products") to the database, the retrieval range can be accurately narrowed. For example, in a customer service system, when a user's question involves the "return and exchange policy", the Agent automatically associates with the "after-sales service database" and preferentially calls its data, improving the response speed.

[0065] Support parallel retrieval of databases with different priorities. For example, high-priority databases (such as core transaction data) use real-time response, and low-priority databases (such as historical logs) use asynchronous processing to balance the system load.

[0066] Step 120: Vectorize the data in the enterprise database to convert the text-form data into a numerical vector representation.

[0067] It should be noted that by converting the text-form data into a vector representation in a high-dimensional space, it is convenient for subsequent fast retrieval and efficient matching.

[0068] As an example, this step can be specifically:

[0069] Obtain the first slice information through a preset slice acquisition interface;

[0070] Based on the first slice information, split the data in the enterprise database to obtain several split segments;

[0071] Obtain the segmentation granularity, and use the segmentation granularity to further split the split segments to obtain several knowledge fields;

[0072] Convert the text-form knowledge fields into a numerical vector representation.

[0073] Those skilled in the art can understand that by obtaining the slicing rules (such as by paragraph, semantic block, or fixed character length) through a preset interface, long texts can be sliced into logically complete segments.

[0074] Example: In a medical knowledge base, a clinical guideline document is sliced into independent segments such as "indications", "medication specifications", and "contraindications" according to the "chapter title" to ensure that each segment carries single-topic information.

[0075] Dynamically define the segmentation granularity (such as sentence level, phrase level) based on business requirements to refine the extraction of knowledge fields.

[0076] Example: After the financial contract text is first segmented by "clause items", it is then segmented again to extract key fields (such as "interest rate = 5%" and "liquidated damages clause = Article 3.2"), which facilitates subsequent precise matching.

[0077] Through hierarchical segmentation, redundant content (such as irrelevant descriptions) is removed, focusing on the core knowledge units.

[0078] The finer the granularity, the more precise the semantic representation after vectorization, reducing the risk of "false recall" (such as confusing "Apple Inc." with "fruit apple").

[0079] Use pre-trained models (such as BERT, Sentence-BERT) to transform the text into dense vectors, capturing deep semantic features.

[0080] Example: In the e-commerce scenario, after vectorizing the product description "waterproof Bluetooth headset", the vector similarity with "sports sweatproof wireless earbuds" reaches 0.92, achieving semantic matching across keywords.

[0081] Traditional keyword retrieval cannot recognize synonyms or context differences (such as "Java" referring to a programming language or coffee beans), while vector similarity can distinguish semantics.

[0082] Case: In the customer service system, the user's question "how to reset the password" can be associated with multiple documents such as "Account Security - Password Modification Process" and "Solutions for Login Abnormalities", even if the word "reset" is not included.

[0083] After vectorization, through the approximate nearest neighbor search (ANN) algorithm (such as FAISS, HNSW), the retrieval time can be reduced from minutes to milliseconds, especially suitable for scenarios with billions of data volumes.

[0084] Step 130: Through the user interface, obtain the user demand data, then extract the retrieval constraint prompt words, and then determine the preset retrieval priorities corresponding to each enterprise database.

[0085] This step can be specifically:

[0086] Through the user interface, obtain the user demand data;

[0087] Input the user demand data into the large model, extract the key features in the user demand data, and then determine whether the key features belong to the preset trigger conditions;

[0088] When it belongs to the preset trigger conditions, extract the retrieval constraint prompt words from the user demand data;

[0089] Based on the correspondence between enterprise databases, retrieval constraint prompt words, and preset retrieval priorities, determine the preset retrieval priorities corresponding to each enterprise database.

[0090] Those skilled in the art can understand that in this step, the original requirements (such as "I want to find customers in North China with sales exceeding 1 million in the last three months") are obtained through an interactive interface (such as a chat window, form). The large language model (such as GPT, ERNIE) is used to parse the requirement text and extract key features (such as the time range "the last three months", the region "North China", and the business indicator "sales > 1 million"). The key features are compared with preset rules (such as "region is required" and "amount threshold triggers the risk control library") to screen out valid constraint words.

[0091] ‌Example‌: The user requirement "need to handle customer complaints" may imply different scenarios (such as "product quality problems" or "logistics delays"). The "complaint type" feature is extracted through the large model and associated with the "after-sales problem library" (priority 8) or the "logistics tracking library" (priority 6).

[0092] ‌Example‌: When the user asks "which suppliers meet the ESG standards", the large model recognizes "ESG" as the environmental, social, and governance standard, extracts the constraint word "ESG rating ≥ A", and preferentially retrieves the "supplier compliance library" (priority 9).

[0093] In this step, the association relationship between the enterprise database and the constraint word can be predefined (such as "sales > 1 million" is bound to the "high-value customer library" with a priority of 10). When multiple constraint words are triggered, the priority is accumulated according to the weight (such as "North China region + sales > 1 million" increases the priority of the "high-value customer library" from 10 to 12).

[0094] In this step, business rules (such as "amount > threshold" and "contains sensitive words") can be preset and automatically matched through regular expressions or classification models. When the trigger condition is met, the default priority is temporarily overwritten (such as the daily priority is 5, and it rises to 12 after the risk control condition is triggered).

[0095] Step 140: Calculate the vector representation with the highest similarity to the retrieval constraint prompt words based on the adjustment weight corresponding to the preset retrieval priority; perform data analysis on the retrieval constraint prompt words to determine the corresponding guiding data; input the guiding data and the vector representation into the large model to obtain the output data.

[0096] Among them, the guiding data at least includes: role positioning guiding words, ability positioning guiding words, task objectives, and output requirements;

[0097] In some embodiments, calculating the vector representation with the highest similarity to the retrieval constraint prompt words based on the adjustment weight corresponding to the preset retrieval priority specifically includes:

[0098] Calculate the initial similarity between the retrieval constraint prompt and the vector representation through the language similarity algorithm;

[0099] Multiply the adjusted weight by the initial similarity to obtain the final similarity, and then obtain the vector representation with the highest similarity.

[0100] Before performing data analysis on the retrieval constraint prompt to determine the corresponding guiding data, the method further includes:

[0101] Configure the correspondence between the retrieval constraint prompt and the guiding data.

[0102] Those skilled in the art can understand that this step can use language similarity algorithms (such as cosine similarity, BM25) to calculate the semantic matching degree between the retrieval constraint prompt and the vector representation. Multiply the adjusted weight corresponding to the preset retrieval priority (such as when the priority of the enterprise database is 10, the weight = 1.2) by the initial similarity to dynamically correct the matching result.

[0103] Case: In the financial risk control scenario, if the priority of the "high-risk transaction rule library" is 15 (weight = 1.5), even if its initial similarity is 0.7, it is increased to 1.05 after weighting, and the risk control interception logic is triggered preferentially.

[0104] Case: Retrieve the "customer demand library" (weight = 1.0) and the "compliance clause library" (weight = 1.3) simultaneously to ensure that the generated content meets both customer demands and regulatory requirements.

[0105] This step restricts the language style and knowledge boundary by defining the identity of the generation subject (such as "professional financial advisor", "medical expert"). Clearly define structured requirements such as the generation format (such as JSON, table), content length, and stop words.

[0106] Case: When configuring "role = legal advisor", the generated text automatically quotes relevant laws (such as Article XXX of the Civil Code) to avoid insufficient professionalism caused by general descriptions.

[0107] Case: Limit the generated content to only be based on the retrieved vector data through the "output requirements" to reduce the probability of fictional information (such as descriptions of drug side effects without retrieval).

[0108] This step can pre-define the weight coefficients of different priority databases (such as the weight of the core business library = 1.5, the weight of the auxiliary library = 0.8). Establish the association rules between the retrieval constraint words and the guiding data (such as "complaint handling" is mapped to "role = customer service specialist", "task goal = resolve within 24 hours").

[0109] Case: When the user inputs "recommend financial products", a solution that matches the user's risk preference is generated based on the priority weight (1.2) of the "product library" and the guiding term "risk level = stable".

[0110] Based on the above description, the present application discloses an enterprise data retrieval method based on a large model. By presetting the retrieval priority, it is ensured that high-value databases (such as core business systems) have a higher weight in the retrieval, avoiding interference from low-priority databases (such as log backups) on key results. Retrieving high-weight databases first reduces unnecessary computational overhead and improves the system response speed. The combination of similarity calculation and priority weight makes the retrieval results more in line with the actual business needs. For example, financial data is preferentially displayed instead of non-critical documents.

[0111] Existing solutions lack active guidance for user needs and are prone to retrieval deviations due to vague requirements. By role positioning (such as "financial analyst") and ability positioning (such as "data visualization"), the user's identity and skill background are clarified, reducing semantic ambiguity. The introduction of task objectives (such as "generate quarterly reports") and output requirements (such as "Excel format") transforms vague requirements into executable retrieval conditions. Differentiated guidance strategies are generated for different user roles (such as sales, R & D) to improve the adaptability of the retrieval experience.

[0112] Existing solutions have insufficient capabilities for processing unstructured data and are difficult to unify the semantic space. Vector representation breaks through the limitations of text form and supports semantic matching of cross-modal data (such as documents, tables, logs). Numerical vectors can be directly input into the large model for reasoning. For example, relevant content in different databases can be associated through vector similarity calculation. Support for multi-condition combined queries (such as "find product defect records in customer complaints") improves the flexibility of retrieval.

[0113] Existing solutions rely on a single retrieval model and are difficult to handle complex tasks. The large model combines the task objective (such as "risk assessment report") in the guiding data and the vector retrieval results to generate structured and highly readable outputs (such as reports, charts). The large model can correct potential biases in the retrieval results (such as similarity calculation errors) and supplement missing information through context reasoning. Support for generating diverse output forms (such as text summaries, visual charts) according to user needs improves the practicality of the results.

[0114] In addition, the present application Figure 2 is an enterprise data retrieval system based on a large model provided by an embodiment of the present application. As Figure 2 shown, the system provided by the embodiment of the present application mainly includes:

[0115] The RAG module 210 is used to create a number of enterprise databases and configure the correspondence between enterprise databases, retrieval constraint prompt words, and preset retrieval priorities; perform vectorization processing on the data in the enterprise databases to convert the data in text form into a numerical vector representation.

[0116] The RAG module 210 includes a revision unit,

[0117] which is used to modify the data in the enterprise database through the backend interface and / or modify the correspondence between the enterprise database, retrieval constraint prompt words, and preset retrieval priorities.

[0118] The user interaction module 220 is used to obtain user requirement data through the user interaction interface;

[0119] The large model module 230 is used to extract retrieval constraint prompt words, and then determine the preset retrieval priorities corresponding to each enterprise database; calculate the vector representation with the highest similarity to the retrieval constraint prompt words with the adjustment weights corresponding to the preset retrieval priorities; perform data analysis on the retrieval constraint prompt words to determine the corresponding guiding data; wherein, the guiding data at least includes: role positioning prompt words, ability positioning prompt words, task objectives, and output requirements; input the guiding data and the vector representation into the large model to obtain output data.

[0120] The large model module 230 includes a determination unit,

[0121] which is used to extract the key features in the user requirement data, and then determine whether the key features belong to the preset trigger conditions;

[0122] when belonging to the preset trigger conditions, extract the retrieval constraint prompt words from the user requirement data;

[0123] Based on the correspondence between the enterprise database, retrieval constraint prompt words, and preset retrieval priorities, determine the preset retrieval priorities corresponding to each enterprise database.

[0124] In addition, the embodiment of the present application also provides a non-volatile computer storage medium, on which executable instructions are stored, and when the executable instructions are executed, the above-mentioned enterprise data retrieval method based on a large model is implemented.

[0125] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for enterprise data retrieval based on a large model, characterized in that: The method comprises: Create several enterprise databases, and configure the corresponding relationships among the enterprise databases, search constraint prompt words, and preset search priorities; Vectorize the data in the enterprise database and convert the text data into numerical vector representation; Through the user interaction interface, user demand data is obtained, and then search constraint prompt words are extracted, and then the preset search priority corresponding to each enterprise database is determined; Using the adjustment weight corresponding to the preset search priority, calculate the vector representation with the highest similarity to the search constraint prompt word; Performing data analysis on the search constraint prompt words to determine corresponding guide data; wherein the guide data at least includes: role positioning guide words, ability positioning guide words, task objectives and output requirements; The bootstrap data and vector representation are input into the large model to obtain the output data.

2. The enterprise data retrieval method based on a large model according to claim 1 is characterized in that: After creating a number of enterprise databases and configuring the preset search priorities of the enterprise databases, the method further includes: Modify the data in the enterprise database through the back-end interface, and / or modify the correspondence between the enterprise database, the search constraint prompt words and the preset search priority.

3. The enterprise data retrieval method based on a large model according to claim 1 is characterized in that: Vectorize the data in the enterprise database and convert the text data into numerical vector representation, including: Obtain the first slice information through the preset slice acquisition interface; Based on the first slicing information, the data in the enterprise database is segmented to obtain a number of segmented fragments; Obtain segmentation granularity, and use the segmentation granularity to segment the segmented fragments again to obtain several knowledge fields; Convert textual knowledge fields into numerical vector representations.

4. The enterprise data retrieval method based on a large model according to claim 1 is characterized in that: Through the user interaction interface, user demand data is obtained, and then the search constraint prompt words are extracted, and then the preset search priority corresponding to each enterprise database is determined, including: Obtain user demand data through the user interaction interface; Input user demand data into the big model, extract key features from the user demand data, and then determine whether the key features belong to the preset trigger conditions; When the preset trigger condition is met, the search constraint prompt words are extracted from the user demand data; Based on the correspondence between the enterprise databases, the search constraint prompt words and the preset search priorities, the preset search priorities corresponding to the various enterprise databases are determined.

5. The enterprise data retrieval method based on a large model according to claim 1 is characterized in that: The vector representation with the highest similarity to the search constraint prompt word is calculated using the adjustment weight corresponding to the preset search priority, including: Calculate the initial similarity between the retrieval constraint prompt word and the vector representation through the language similarity algorithm; By adjusting the weight and multiplying it with the initial similarity, the final similarity is obtained, and then the vector representation with the highest similarity is obtained.

6. The enterprise data retrieval method based on a large model according to claim 1 is characterized in that: Before performing data analysis on the search constraint prompt words to determine corresponding guide data, the method further includes: Configure the correspondence between search constraint hint words and guide data.

7. An enterprise data retrieval system based on a large model, characterized in that: The system comprises: RAG module is used to create several enterprise databases and configure the corresponding relationship between enterprise databases, search constraint prompt words and preset search priorities; vectorize the data in the enterprise database and convert the text data into numerical vector representation; A user interaction module is used to obtain user demand data through a user interaction interface; The large model module is used to extract search constraint prompt words, and then determine the preset search priority corresponding to each enterprise database; calculate the vector representation with the highest similarity to the search constraint prompt word with the adjustment weight corresponding to the preset search priority; perform data analysis on the search constraint prompt word to determine the corresponding guide data; wherein the guide data at least includes: role positioning guide words, ability positioning guide words, task objectives and output requirements; input the guide data and the vector representation into the large model to obtain output data.

8. The enterprise data retrieval system based on a large model according to claim 7, characterized in that: The RAG module includes revision units, Used to modify data in the enterprise database through the back-end interface, and / or modify the correspondence between the enterprise database, search constraint prompt words and preset search priorities.

9. The enterprise data retrieval system based on a large model according to claim 7, characterized in that: The large model module includes a determination unit, Used to extract key features from user demand data and then determine whether the key features belong to the preset trigger conditions; When the preset trigger condition is met, the search constraint prompt words are extracted from the user demand data; Based on the correspondence between the enterprise databases, the search constraint prompt words and the preset search priorities, the preset search priorities corresponding to the various enterprise databases are determined.

10. A non-volatile computer storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed, the method for enterprise data retrieval based on a large model as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Large model knowledge question and answer optimization method based on semantic enhanced knowledge graph

    CN118939779A

  • AI large model enterprise knowledge base system based on cue words

    CN119106143A